Paper deep dive
SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia
Ri Chi Ng, Aditi Kumaresan, Yujia Hu, Roy Ka-Wei Lee
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 3/22/2026, 5:29:49 AM
Summary
SEAHateCheck is a pioneering functional testing framework and dataset designed to evaluate hate speech detection models in low-resource Southeast Asian languages, specifically Indonesian, Tagalog, Thai, and Vietnamese. By leveraging HateCheck's methodology, local expert validation, and LLM-augmented generation, the study provides a robust benchmark for identifying model weaknesses in implicit hate detection, slang-based expressions, and tonal linguistic nuances.
Entities (8)
Relation Signals (4)
SEAHateCheck → builton → HateCheck
confidence 100% · Building on HateCheck's functional testing framework
SEAHateCheck → coverslanguages → Indonesian, Tagalog, Thai, Vietnamese
confidence 100% · SEAHateCheck, a pioneering dataset tailored to Indonesia, Thailand, the Philippines, and Vietnam, covering Indonesian, Tagalog, Thai, and Vietnamese.
SEAHateCheck → refines → SGHateCheck
confidence 100% · refining SGHateCheck's methods
SEA-Lionv2.1 → usedtogenerate → SEAHateCheck
confidence 95% · Prompts were tested on Ministral-8B-Instruct-2410 and SEA-Lionv2.1, with SEA-Lionv2.1 selected for lower safety guardrail rejections.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Hate speech detection relies heavily on linguistic resources, which are primarily available in high-resource languages such as English and Chinese, creating barriers for researchers and platforms developing tools for low-resource languages in Southeast Asia, where diverse socio-linguistic contexts complicate online hate moderation. To address this, we introduce SEAHateCheck, a pioneering dataset tailored to Indonesia, Thailand, the Philippines, and Vietnam, covering Indonesian, Tagalog, Thai, and Vietnamese. Building on HateCheck's functional testing framework and refining SGHateCheck's methods, SEAHateCheck provides culturally relevant test cases, augmented by large language models and validated by local experts for accuracy. Experiments with state-of-the-art and multilingual models revealed limitations in detecting hate speech in specific low-resource languages. In particular, Tagalog test cases showed the lowest model accuracy, likely due to linguistic complexity and limited training data. In contrast, slang-based functional tests proved the hardest, as models struggled with culturally nuanced expressions. The diagnostic insights of SEAHateCheck further exposed model weaknesses in implicit hate detection and models' struggles with counter-speech expression. As the first functional test suite for these Southeast Asian languages, this work equips researchers with a robust benchmark, advancing the development of practical, culturally attuned hate speech detection tools for inclusive online content moderation.
Tags
Links
- Source: https://arxiv.org/abs/2603.16070v1
- Canonical: https://arxiv.org/abs/2603.16070v1
Trouble viewing inline? Open PDF directly →
Full Text
236,748 characters extracted from source content.
Expand or collapse full text
SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia RI CHI NG ∗ , Singapore University of Technology and Design, Singapore ADITI KUMARESAN ∗ , Singapore University of Technology and Design, Singapore YUJIA HU, Singapore University of Technology and Design, Singapore ROY KA-WEI LEE, Singapore University of Technology and Design, Singapore Hate speech detection relies heavily on linguistic resources, which are primarily available in high-resource languages such as English and Chinese, creating barriers for researchers and platforms developing tools for low-resource languages in Southeast Asia, where diverse socio-linguistic contexts complicate online hate moderation. To address this, we introduce SEAHateCheck, a pioneering dataset tailored to Indonesia, Thailand, the Philippines, and Vietnam, covering Indonesian, Tagalog, Thai, and Vietnamese. Building on HateCheck’s functional testing framework and refining SGHateCheck’s methods, SEAHateCheck provides culturally relevant test cases, augmented by large language models and validated by local experts for accuracy. Experiments with state-of-the-art and multilingual models revealed limitations in detecting hate speech in specific low-resource languages. In particular, Tagalog test cases showed the lowest model accuracy, likely due to linguistic complexity and limited training data. In contrast, slang-based functional tests proved the hardest, as models struggled with culturally nuanced expressions. The diagnostic insights of SEAHateCheck further exposed model weaknesses in implicit hate detection and models’ struggles with counter-speech expression. As the first functional test suite for these Southeast Asian languages, this work equips researchers with a robust benchmark, advancing the development of practical, culturally attuned hate speech detection tools for inclusive online content moderation. CCS Concepts:• Computing methodologies→ Natural language processing; Language resources. Additional Key Words and Phrases: Hate Speech, Low-Resource Languages, Benchmarks ACM Reference Format: Ri Chi Ng, Aditi Kumaresan, Yujia Hu, and Roy Ka-Wei Lee. 2026. SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia. 1, 1 (March 2026), 95 pages. https://doi.org/X.X 1 Introduction Online hate speech in Southeast Asia (SEA) fuels discrimination, division, and harm targeted at vulnerable communities. However, detection models struggle to address this crisis in low-resource languages like Indonesian, Malay, Tagalog, Thai, and Vietnamese. These languages, encompassing tonal systems (Thai, Vietnamese) and script-based orthographies (Indonesian, Malay, Tagalog), are underrepresented in hate speech datasets, which are predominantly trained on ∗ Both authors contributed equally to this research. Authors’ Contact Information: Ri Chi Ng, richi_ng@sutd.edu.sg, Singapore University of Technology and Design, Singapore, Singapore; Aditi Kumaresan, aditi_kumaresan@sutd.edu.sg, Singapore University of Technology and Design, Singapore, Singapore; Yujia Hu, yujia_hu@sutd.edu.sg, Singapore University of Technology and Design, Singapore, Singapore; Roy Ka-Wei Lee, roy_lee@sutd.edu.sg, Singapore University of Technology and Design, Singapore, Singapore. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. Manuscript submitted to ACM Manuscript submitted to ACM1 arXiv:2603.16070v1 [cs.CL] 17 Mar 2026 2Ng et al. high-resource languages such as English and Mandarin [51]. This bias exacerbates the challenges in capturing the sociolinguistic complexity of SEA, where culturally nuanced expressions - slang, implicit insults, and coded hate - permeate online discourse. As social networks amplifies hate, targeting marginalized groups such as ethnic minorities and religious communities (e.g., those identified in local legislation, Table 11), social networks are not equipped to moderate content effectively. The absence of robust detection tools not only undermines online safety but also risks deepening social tensions in a region marked by diverse histories and identities. Urgent action is needed to develop culturally attuned hate speech detection systems that reflect SEA’s unique linguistic and cultural landscape. Functional testing frameworks have emerged to address limitations in traditional hate speech evaluation, which relies on held-out test sets prone to biases and gaps. HateCheck [41] introduced targeted test cases to assess model performance in English, focusing on real-world scenarios like negation and identity-based hate. Multilingual HateCheck (MHC) [40] extended this approach to other high-resource languages, while SGHateCheck [30] adapted it for Singapore’s multilingual context, incorporating local slang and cultural references. Despite these advances, these frameworks are inadequate for the broader low-resource languages of SEA, which require customized test cases to address tonal phonetics, script diversity, and region-specific hate speech patterns (e.g., implicit hate in Thai proverbs or Vietnamese online forums). This gap leaves researchers and platforms without the tools to comprehensively evaluate hate speech detection in the diverse settings of the SEA. To bridge this critical gap, we introduce SEAHateCheck, the first functional test suite designed to evaluate hate speech detection models across SEA. It builds on the dataset created in SGHateCheck and covers the sociocultural context of Indonesia, Malaysia, the Philippines, Singapore, Thailand, and Vietnam, and covers a wide array of languages including Indonesian, Malay, Mandarin, Singlish, Tagalog, Tamil, Thai, and Vietnamese 1 . Building on HateCheck’s robust framework and refining SGHateCheck’s localization techniques, SEAHateCheck delivers a comprehensive set of culturally relevant test cases, addressing slang, implicit hate, and vulnerable groups identified through local expertise (Section 2.1). By integrating large language models (LLMs) for test case generation, native speakers for accurate translations, and local experts for cultural validation, SEAHateCheck sets a new standard for hate speech evaluation in low-resource settings. As a diagnostic benchmark, it empowers researchers to assess model performance systematically, fostering the development of inclusive and effective hate speech detection tools tailored to SEA’s unique needs. SEAHateCheck’s contributions extend beyond its pioneering dataset, offering actionable insights from rigorous evaluation of state-of-the-art LLMs. Our findings reveal critical model weaknesses, such as lower accuracy in Vietnamese test cases, which is likely due to the language’s tonal complexity and limited training data, as well as struggles with slang-based tests that require cultural nuance (e.g., region-specific colloquialisms). These insights guide developers to prioritize enhanced training for tonal languages and context-aware algorithms, addressing gaps in current models. For platforms, SEAHateCheck informs moderation strategies to protect marginalized groups better, aligning with local legislation on protected categories. Its diagnostic capabilities further highlight deficiencies in detecting implicit hate, enabling targeted improvements in model robustness. By providing a scalable, culturally grounded benchmark, SEAHateCheck transforms hate speech detection, safeguarding SEA’s diverse communities and paving the way for equitable online moderation globally. 1 Dataset available at https://github.com/Social-AI-Studio/SEAHateCheck Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia3 Protected Categories Indonesia Malaysia the Philippines Singapore Thailand Vietnam ReligionYesYesYesYesYesYes Ethnicity/Race/OriginYesYesYesYesYesYes DisabilitiesNoYesYesYesYesYes Gender/OrientationYesYesYesYesYesYes AgeNoNoNoYesYesYes Vulnerable WorkersNoNoNoNoYesNo People Living with HIVNoNoYesNoNoYes Table 1. Protected categories represented for each country in SEAHateCheck. 2 Constructing SEAHateCheck Dataset 2.1 Defining Hate Speech SEAHateCheck adopts HateCheck’s definition of hate speech as “abuse aimed at a protected group or its members for belonging to that group” [41]. Following SGHateCheck’s approach, local cultural experts (two per country, with backgrounds in sociology and linguistics) consulted legislation and used it as guidance to suggest protected categories. Within each protected category is a protected target (e.g., Hindus (protected target) for Religion(protected group)). Table 1 details these categories, and Appendix Table 11 lists legislative sources. All countries share four common categories—Religion, Ethnicity/Race/Origin, Disabilities, and Gender—with additional categories (e.g., Age, People Living with HIV) varying by country, ensuring cultural relevance. The legislative and regulatory sources consulted to define these protected categories are listed in Table 11. 2.2 Defining Functional Tests In SEAHateCheck, a functional test is defined as a targeted evaluation of a hate speech detection model’s ability to correctly classify short text statements as hateful or non-hateful, following the diagnostic framework introduced by HateCheck (Röttger et al., 2021). Each test targets a specific functionality, such as distinguishing hate speech containing profanity from non-hateful expressions with similar lexical features. For instance, a test case in Tagalog may assess hateful profanity directed at a protected group (e.g., “Tangina, ang hirap nun,” targeting a group), contrasted with a non-hateful, colloquial use of profanity in Tagalog (e.g., “Bakit ba hindi tumitigil ang pag-iyak ng mga sanggol sa eroplano?”). Tests are designed to be fine-grained, contrastive, and culturally relevant, enabling models to discern nuanced language use across diverse Southeast Asian contexts. To facilitate systematic evaluation, we organize tests into thematic categories, such as explicit hate, implicit hate, and non-hateful contrasts, aligning with the sociolinguistic use in Indonesia, Malaysia, the Philippines, Singapore, Thailand, and Vietnam. This structure enhances diagnostic insights into model performance, revealing whether models rely on superficial cues or capture context-specific hate speech patterns. 2.3 Selecting Functional Tests SEAHateCheck’s functional tests were selected to align with the HateCheck framework [41], adapting its methodology to the sociolinguistic contexts of Indonesia, the Philippines, Thailand, and Vietnam. Following HateCheck’s approach, which integrates interviews with NGO workers and a review of hate speech research, we incorporated country-specific elements through consultations with local experts in sociology and linguistics. This ensures that our tests are culturally attuned, enhancing their relevance for evaluating hate speech detection models in each country’s unique context. Manuscript submitted to ACM 4Ng et al. All test cases are short text statements, designed to be unambiguously hateful or non-hateful per our hate speech definition. SEAHateCheck comprises up to 27 functional tests per language (22 for Malay, Tamil, Indonesian, Tagalog, Thai, and Vietnamese; 25 for Mandarin; 27 for Singlish), tailored to reflect linguistic and cultural considerations. For instance, we excluded slur homonyms and reclaimed slurs absent in Indonesian, Malay, Mandarin, Tagalog, Tamil, Thai, Singlish, and Vietnamese, and omitted spelling variations to streamline translation. Like HateCheck and MHC, our tests distinguish hate speech from non-hateful content with similar lexical features but clear non-hateful intent, enabling nuanced evaluation across diverse expressions. A table summarising the functional tests, together with examples in represented languages and the original English templates is shown in Fig 5 and Fig 6. The targets in both tables were replaced with a placeholder TARGET. Distinct Expressions of Hate. SEAHateCheck evaluates varied forms of hate speech, including derogatory remarks (F1–F4) and threats (F5–F6), as well as hate conveyed through slurs (F7) and profanity (F8). It assesses hate expressed via pronoun references (F10–F11), negation (F12), and varied phrasings, such as questions and opinions (F14–F15). For Indonesian, Tagalog, and Vietnamese, tests include spelling variations like omissions or leet speak (F23–F34), broadening the evaluative scope to capture region-specific linguistic patterns. Contrastive Non-Hate. To ensure robust model evaluation, SEAHateCheck includes non-hateful content, such as profanity used without malice (F9), negation (F13), and benign references to protected groups (F16–F17). It also examines contexts where hate speech is quoted or countered, particularly in counter-speech scenarios that neutralize hate (F18–F19). Additionally, tests differentiate content targeting non-protected entities, such as objects (F20–F22), ensuring clear distinctions between hateful and non-hateful expressions. Text Obfuscations. For Singlish and Mandarin, there are additional functional tests (F23 to F34), where the texts were methodically obfuscated in different ways. 2.4 Translating Templates To adapt HateCheck’s functional test templates [41] for Indonesian, Tagalog, Thai, and Vietnamese, we employed human translators supported by LLMs, producing 655 templates per language. These templates cover 22 functional tests across protected categories, ensuring cultural and linguistic accuracy for low-resource Southeast Asian languages [37]. The translation process differed by language: Indonesian followed SGHateCheck’s protocol [30] (as language experts were available when SGHateCheck was made), while Tagalog, Thai, and Vietnamese used a three-stage approach. For the Indonesian data, we began by fine-tuning GPT-3.5 [3] on a small set of 27 human-translated templates from SGHateCheck so that the model could better handle hate speech contexts. We then used the fine-tuned model to translate all 655 English templates into Indonesian. Two native translators reviewed the generated sentences line by line and edited them to match Indonesian usage, replacing overly literal phrases, removing inappropriate slurs, and adjusting the tone to sound natural. For Tagalog, Thai, and Vietnamese, the same set of 655 English templates was translated using Gemini 1.5 Pro [49] and GPT-4o [32]. Two translators per language selected or edited machine-generated translations, or provided original translations when needed, ensuring sociolinguistic relevance. The process unfolded in three stages: (1) Stage 1: Template Validation: One template per functional test (22 templates) was translated and reviewed with sociolinguistic experts of the respective countries to confirm cultural applicability. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia5 (2)Stage 2: Multi-Shot Translation: Using Stage 1 translations as multi-shot examples (i.e., in-context learning prompts), 100 additional templates were translated and edited, with translators resolving discrepancies through discussion between translators. (3) Stage 3: Scaled Translation: Using randomly selected Stage 2 translations as examples, 611 templates were translated, completing the set. Discrepancies were also resolved here through discussion between translators. We document language varieties and annotator localities to ground translation choices in Appendix A.2 - A.3. A detailed breakdown of the templates and generated test cases across functional tests and languages is provided in Table 14. 2.5 Generating and Validating Gold Label Test Cases Gold label test cases in SEAHateCheck are manually validated instances derived from translated templates, designed to evaluate hate speech detection models across Indonesian, Tagalog, Thai, and Vietnamese. This subsection outlines the pipeline for generating and validating these test cases, ensuring cultural relevance and linguistic accuracy [41]. Generating Test Cases. Using 655 templates per language, we generated test cases by substituting placeholders for protected groups (e.g., IDENTITY_P) and slurs (e.g., SLUR_P). We summarize the target groups covered by the training corpora for each language in Table 5. Three experts per language identified up to six protected categories per country (See Table 1) (Protected categories are further broken down into protected groups, e.g., Hindus, Muslims for Religion; ethnic Chinese for Ethnicity) and 10–15 common slurs, based on local demographics and online discourse. For example, the Indonesian template Kita bakal terus serang lo IDENTITY_P sampe lu pergi” becomes Kita bakal terus serang lo Muslim sampe lu pergi,” while a non-hateful Tagalog template Kami ay may respeto sa IDENTITY_P” yields Kami ay may respeto sa Muslim.” Each template was instantiated 5–10 times per group, producing 31,172 test cases across four languages (7,793.5 per language on average; Table 2). Of these, 21,187 were labeled hateful and 9,985 non-hateful based on template sentiment [41]. Test cases averaged 10.4 words (excluding Thai due to lack of word separators) and 50.3 Validating Test Cases. To ensure quality, 16,415 test cases (approximately 4,104 per language) were annotated by 12 native speakers (3 per language) with linguistics training, each reviewing cases in triplicate. Annotators labeled sentiment (Hateful,” Non-hateful,” or Nonsensical”) per Section 2.1 and flagged cases for quality issues (unnatural phrasing or context dependence). Training on 50 sample cases ensured consistency. High-quality test cases required unanimous sentiment agreement, alignment with template sentiment, and no quality flags. Of 16,415 cases, ̃ 5% (820) were labeled Nonsensical” due to translation errors or cultural mismatches and excluded. After filtering, 13,579 high-quality test cases were retained for benchmarking (Section 3), with a mean high-quality rate of 0.83 (proportion of cases meeting all criteria; Table 2). Inter-annotator agreement (Fleiss’ kappa = 0.85) indicates high reliability, detailed in Appendix A.5. Comparison with SGHateCheck. Like SGHateCheck [30], SEAHateCheck uses a shared pool of 655 templates to generate test cases, which ensures comparable functional coverage across Southeast Asian languages. SEAHateCheck focuses on four low resource languages, yielding 13,579 validated gold label test cases, whereas SGHateCheck provides roughly 11,000 validated cases for Malay, Singlish, Tamil and Chinese. Both datasets rely on native annotators, but SEAHateCheck applies a stricter unanimous agreement criterion, which results in a slightly lower high quality rate. 2.6 Generating and Validating Silver Label Test Cases Silver label test cases in SEAHateCheck are LLM-generated instances that enhance the localization and scale of hate speech detection for Indonesian, Tagalog, Thai, and Vietnamese, complementing template-based gold test cases (Section Manuscript submitted to ACM 6Ng et al. 2.5). This subsection details their motivation, generation, validation, limitations, and comparison with SGHateCheck [30]. Motivation. Silver test cases address limitations of gold test cases, which rely on manually translated templates (Section 2.5). First, LLMs enable rapid scaling, producing 19,802 test cases compared to 13,579 gold, covering diverse hate speech scenarios. Second, they capture colloquial expressions (e.g., “bajingan” in Indonesian) missed by templates, enhancing realism for low-resource languages. Third, they reduce annotation costs, validating 400 cases vs. 16,415 for gold. Finally, their variability tests model robustness against naturalistic inputs, providing complementary diagnostics [37]. Generating Test Cases. We used 13,579 high-quality gold test cases (Section 2.5), grouped by 22 functional tests and 6 protected groups (e.g., Muslims, ethnic Chinese), yielding ~100 groups per language. Multi-shot prompts with 3–5 gold test cases were designed with three native speakers per language over two iterations to ensure casual, localized outputs (e.g., Use slang like ‘bajingan’ in Indonesian). Prompts were tested on Ministral-8B-Instruct-2410 [1] and SEA-Lionv2.1 [46], with SEA-Lionv2.1 selected for lower safety guardrail rejections. Generation produced 10 test cases per group, yielding 19,802 test cases (4,950.5 per language; Table 2). Of these, 14,145 were intended as hateful and 5,657 non-hateful, based on prompting gold test cases’ sentiment. Test cases averaged 15.6 words (excluding Thai) and 74.2 characters, ~50% longer than gold due to LLM verbosity (e.g., qualifiers like sangat”). Validating Test Cases. To assess quality, 100 test cases per language (400 total) were annotated by 12 native speakers (3 per language) with linguistics training, each reviewed in triplicate. Annotators labeled sentiment (Hateful,” Non-hateful,” “Nonsensical”), flagged unnatural phrasing or context dependence (per Section 2.5), and verified target group and functional test matching. Training on 50 sample cases ensured consistency. Fifteen positive controls (same group/test) and 15 negative controls (same group, different test) confirmed annotator accuracy and test specificity. High-quality test cases required unanimous sentiment agreement, no quality flags, and matching target/function. The high-quality rate (0.72 mean; Table 2) reflects 72% of test cases meeting all criteria, vs. 83% for gold, due to LLM variability. Inter-annotator agreement (Fleiss’ kappa = 0.80) is reliable but lower than gold (0.85), detailed in Appendix A.5. A Quality Score (0–5) awarded 1 point each for correct sentiment, naturalness, context independence, target, and function. Table 3 shows silver scores (e.g., 0.74 for sentiment) are lower than gold (0.90). Tamil’s concise silver test cases (6.7 vs. 7.2 words; Table 2) reflect LLM constraints in agglutinative languages. Code-switching. Our template-based Gold cases were translated with a preference for predominantly monolingual realizations to preserve controlled functional contrasts. We did not intentionally design code-mixed (“Taglish”, “In- doglish”) test cases as a separate condition. Code-switching is prevalent in SEA online discourse and has been studied as a distinct evaluation setting for LLM translation, suggesting the need for dedicated code-mixed test suites [21]. We view systematic code-switching as an important extension for Southeast Asia, where mixing is a frequent evasion tactic. We therefore include code-mixed functional tests as future work. Limitations. Silver test cases face several challenges. First, their lower high-quality rate (0.72 vs. 0.83) and quality scores (Table 3) indicate reduced naturalness and reliability, as LLMs introduce verbosity or errors. Second, inconsistent target group and functional test alignment (e.g., scores of 0.70–0.72 vs. 0.91–0.92) reduces diagnostic precision, as LLMs may deviate from prompts. Third, despite native speaker input, LLMs struggle with cultural nuances in low-resource languages (e.g., Thai’s tonal complexity), leading to unnatural outputs. Finally, safety guardrails limit the generation of certain hateful content, potentially skewing the dataset. These issues, coupled with dependence on gold test cases’ biases, require rigorous validation, partially offsetting cost savings. Prior work also shows cross-lingual transfer can Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia7 SEAHateCheckSGHateCheck [30] MetricID TL THVI MeanMS SG TA ZH ‡ Mean Gold Label # Test cases819087518488103198937NANANANANA # Hateful template 55115902568169986023NANANANANA # Non-hateful template26792849280733212914NANANANANA Avg. words8.39.6NA13.210.4 † 9.58.57.215.610.2 Avg. characters50.055.238.857.250.358.945.662.415.645.6 Gold Validation Total Annotated35794072395248124103.822532974285128482731.5 High Quality Rate0.920.850.870.910.890.910.910.900.820.88 Silver Label # Test cases45054870479756304950.527593623259135883140.3 # Hateful template31713477347540223536.319602824184828572372.3 # Non-hateful template13341393132216081414.3799799743731768.0 Avg. words 14.314.3NA18.315.6 † 11.812.26.721.613.1 Avg. characters90.382.843.580.274.372.967.460.121.655.5 Silver Validation Total Annotated100100100100100100100100100100 High Quality Rate0.720.610.750.800.720.680.860.680.760.74 Table 2. Summary of SEAHateCheck and SGHateCheck test instances by language and functional test category, including the number of gold and silver cases and their hateful vs. non-hateful splits. Column headers use language abbreviations: ID=Indonesian, TL=Tagalog, TH=Thai, VI=Vietnamese, MS=Malay, SG=Singlish, TA=Tamil, ZH=Mandarin. High quality rate denotes the proportion of high quality annotations, defined in §2.5 and §2.6 for Silver Label test cases. NA marks metrics not reported in the source table. † Weighted average words exclude Thai due to missing word segmentation; Thai cells therefore omit word counts. ‡ For Mandarin (ZH), each character is counted as one word (words = characters). significantly affect hallucination behavior in low-resource settings, motivating conservative filtering and targeted validation for LLM-generated cases [53]. Comparison with SGHateCheck. SGHateCheck [30] generates 12,561 silver test cases for Malay, Singlish, Tamil, and Mandarin, using similar LLM prompting. SEAHateCheck’s 19,802 test cases reflect its low-resource focus, with a slightly lower high-quality rate (0.72 vs. 0.74; Table 2) due to stricter filtering. Quality scores (Table 3) show consistent trends across datasets. These 19,802 test cases, summarized in Table 2, enhance SEAHateCheck’s evaluation of hate speech detection, despite limitations. SGHateCheck metrics are in Table 2. Further details on corpus size, gold and silver generation counts, and the rationale for targeting SEA socio-linguistic contexts are provided in Appendix A.1. Table 2 summarizes the statistics for SEAHateCheck and SGHateCheck side by side, including the number of test cases, the hateful or non hateful balance and the proportion of high quality items. For a more fine-grained view, Table 12 in Appendix reports the number of templates (#TP) and instantiated test cases (#TC) for each functional test and language, making the coverage of explicit hate, implicit hate, and contrastive non-hate tests fully transparent. Notably, SEAHateCheck maintains a comparable Gold high-quality rate (mean 0.89) to SGHateCheck (mean 0.88) while expanding coverage to additional Southeast Asian languages, including tonal languages such as Thai and Vietnamese. This suggests that the translation and expert validation pipeline scales to more typologically diverse settings without materially degrading dataset quality. 3 Benchmarking LLMs on SEAHateCheck We evaluated SEAHateCheck and SGHateCheck across various open-source and closed-source LLMs to assess their effectiveness in detecting HS. The selected models include state-of-the-art (SOTA) architectures and multilingual models fine-tuned to support the majority of languages present in both datasets, including English (representing Singlish), Manuscript submitted to ACM 8Ng et al. DatasetAverage Quality score LanguageLabel SentimentContextNaturalTargetfunctional test Indonesian Gold Label 0.9780.9930.928NANA Silver Label0.8270.9100.7100.8670.470 Tagalog Gold Label 0.9520.9930.998NANA Silver Label0.7970.8500.8200.8670.607 Thai Gold Label0.9600.9870.986NANA Silver Label 0.8530.9770.8870.9130.700 Vietnamese Gold Label0.9720.9940.991NANA Silver Label 0.9100.9830.9870.9470.650 Malay Silver Label 0.7830.9330.8630.9400.440 Singlish0.9200.9730.9570.9770.437 Tamil 0.8230.9070.8270.9070.667 Mandarin0.8600.8630.9000.8470.583 Table 3. Average scores for different annotation fields for each language. Indonesian, Malay, Mandarin, Tagalog, Tamil, Thai, and Vietnamese. These languages will be collectively referred to as the evaluated languages. Our evaluation follows a two-stage approach. First, we assess each model in its default, out-of-the-box (OOTB) configuration to establish a baseline for its intrinsic HS detection capabilities. Second, we fine-tune the models using a curated HS dataset and re-evaluate their performance to measure the impact of domain-specific adaptation. The characteristics of all evaluated models are detailed in Appendix Table 16. To maintain consistency with the original annotation process, we fine-tune and evaluate the models with the same language-specific prompts used during annotation as detailed in Appendix C. 3.1 LLM Fine-tuning The open-source LLMs were further enhanced by fine-tuning with hate speech scraped from social media. To do so, we curated a specialized dataset that captures high-quality, labelled hate speech (HS) observed in the evaluated languages. Detailed characteristics of the training data, including collection methods, are provided in Table 4. Next, we also highlight the observed target groups in the curated hate speech data. Finally, we present a comprehensive breakdown of the training data distribution, including the proportion of hateful vs. non-hateful instances per language and category, in Table 6. We use binary labels (hateful or non-hateful) to perform supervised fine-tuning of the LLMs using Low-Rank Adaptation (LoRA) [18]. The exact fine-tuning and evaluation prompts are included verbatim in Appendix C to ensure reproducibility. 4 Discussion on Gold Label Test Cases 4.1 Overall Results In evaluating the non-finetuned models, performance varied across languages. While most models achieved strong F1 scores for languages like Vietnamese and Indonesian, several open-source models struggled with Tamil and Tagalog. Notably, Deepseek consistently underperformed compared to o3 and Gemini, sometimes yielding worse results than the open-source models. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia9 Dataset Dataset NameLanguage Year Collection Method SEA id-multi-label-hate-speech-and- abusive-language-detection [22] Indonesian2019Twitter Philippine Election-Related Tweets [4]Tagalog2019Twitter HateThaiSent [27]Thai2024— ViHSD [25]Vietnamese2021Facebook, YouTube SG HateM [26]Malay2023— COLDataset [12]Mandarin2022Zhizhu, Weibo HateXplain [28]Singlish 12021Twitter, Gab Waseem and Hovy [52]Singlish 22016— TamilMixSentiment [5]Tamil2020YouTube comments Table 4. Details of the Datasets Used, Including Collection Year and Method Region Language Target Group SEA IndonesianReligion/Creed, Race/Ethnicity, Physical/disability, Gender/Sexual Ori- entation, Other invective/slander TagalogRace, Physical, Sex, Disability, Religion, Class, Quality Thai— VietnameseAimed at all groups/individuals MalayRace, Ethnicity, National Origin, Caste, Sexual Orientation, Gender, Gender identity, Religious Affiliation, Age, Disability, or Serious Disease SG MandarinRace, Religion, Sex, or Sexual Orientation Singlish 1African, Islam, Jewish, LGBTQ, Women, Refugee, Arab, Caucasian, His- panic, Asian Singlish 2Race, Sex TamilNo clear Target Groups Table 5. Target Groups for Different Languages After fine-tuning, a marked improvement in precision was observed across all models, which indicates a reduction in false positives. This was particularly evident in Vietnamese, Tagalog, Tamil, and Singlish, where F1 scores exhibited significant gains. However, fine-tuning led to a decrease in performance for some languages, specifically Malay, Thai and Indonesian. This deterioration can likely be attributed to these datasets’ poor distribution of hateful and non-hateful labels, as highlighted in Table 6. This imbalance may have affected their ability to generalise effectively on high-quality test cases. In Table 7, the non-finetuned results reveal clear stratification among both the model architectures and the languages under consideration. Strong general-purpose systems such as Deepseek, o3 and Gemini achieve the highest F1 scores in most evaluations, with o3 reaching 89.01 in Indonesian and 87.63 in Vietnamese, and Gemini exceeding 80 in various Southeast Asian and Singaporean varieties. Among the nine open-sourcing models, Gemma delivers competitive baselines, often in the mid-to-high seventies, while Sealion and Seagem perform notably well for specific languages. In Manuscript submitted to ACM 10Ng et al. Language Full Training DatasetSampled Training Dataset Not Hateful Hateful TotalNot Hateful Hateful Total SEA Indonesian9,0781,45710,5353,5431,4575,000 Tagalog5,3404,66010,0002,5002,5005,000 Thai3,9622,1156,0772,8852,1155,000 Vietnamese21,4902,55624,0462,5002,5005,000 SG Mandarin13,00312,72325,7262,5002,5005,000 Malay2,4011,5123,9132,4011,5123,913 Singlish 110,6354,74815,3832,5002,5005,000 Singlish 28,8264,00212,8282,5002,5005,000 Tamil22,8827,43430,3162,5002,5005,000 Total97,61741,207138,82423,82920,08443,913 Table 6. Statistics of Hateful and Not Hateful Samples in the Full Dataset and Sampled Subset. Language subsets with poor label distribution are highlighted in red. Model SEASG Indonesian Tagalog Thai Vietnamese Malay Mandarin Singlish Tamil Ministral66.5365.9768.8275.1566.2664.9272.2265.66 Llama3b67.3762.8072.1774.3864.2966.4969.5466.75 Llama8b65.1359.4969.0972.3566.1661.6673.9359.39 Sealion72.0858.3870.5474.4370.8568.2571.6673.42 Seallm70.2756.8570.4874.2865.9266.1868.6454.88 Pangea68.2953.7657.0873.9467.3567.0164.2156.83 Qwen73.8661.8372.9277.2069.4773.0673.5258.33 Gemma80.3677.2576.5583.0977.6677.0374.4678.35 Seagem78.7577.1374.5881.8677.2371.8976.9281.07 Gemini84.5878.5076.4979.8480.2874.7979.8881.09 o389.0182.5980.2187.6385.9483.2883.0885.96 Deepseek74.4765.9667.0377.2776.4467.3677.0167.96 Table 7. F1 scores of different non-finetuned models on SEAHateCheck and SGHateCheck High-Quality Test Cases. contrast, lower-capacity or earlier-generation systems, represented by Ministral and Pangea, display weaker baselines, particularly for Tagalog and Tamil. Vietnamese and Malay tend to yield relatively stronger scores for the highest- capacity models. In contrast, Tagalog and Tamil exhibit broader dispersion that aligns with the greater linguistic and sociolinguistic variability evident in hateful content. Table 8 shows F1 scores after finetuning, demonstrating substantial improvements for most open models on SEA languages, as well as Singlish and Tamil. For instance, Gemma benefits significantly in Thai and Vietnamese and shows further gains in Indonesian and Tagalog. SeaLion, Llama-8B, Pangea, and Qwen also display consistent improvements across several languages. These results indicate that domain-specific supervision effectively enhances recall while maintaining precision on the functional tests. At the same time, the table reveals notable regressions: some models Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia11 Model SEASG Indonesian TagalogThaiVietnameseMalayMandarinSinglishTamil Ministral56.49 (10.04)65.61 (0.36)66.87 (1.95)80.05 (4.90)63.00 (3.26)60.77 (4.15)75.33 (3.11)70.48 (4.82) Llama3b57.91 (9.46)66.21 (3.41)67.18 (4.99)78.28 (3.90)61.62 (2.67)58.13 (8.36)65.73 (3.81)68.96 (2.21) Llama8b71.23 (6.10)67.86 (8.37)76.16 (7.07)79.78 (7.43)71.50 (5.34)63.94 (2.28)78.93 (5.00)65.24 (5.85) Sealion76.24 (4.16)74.04 (15.66)74.26 (3.72)82.17 (7.74)75.04 (4.19)63.94 (4.31)82.27 (10.61)77.64 (4.22) Seallm69.73 (0.54)69.60 (12.75)71.16 (0.68)81.75 (7.47)72.31 (6.39)66.27 (0.09)75.16 (6.52)67.92 (13.04) Pangea65.22 (3.07)67.76 (14.00)65.98 (8.90)83.03 (9.09)67.53 (0.18)63.43 (3.58)73.35 (9.14)69.09 (12.26) Qwen72.52 (1.34)67.51 (5.68)74.85 (1.93)79.96 (2.76)71.58 (2.11)62.33 (10.73)78.19 (4.67)70.05 (11.72) Gemma81.72 (1.36)80.26 (3.01)82.32 (5.77)88.68 (5.59)78.28 (0.62)65.03 (12.00)86.16 (11.70)81.26 (2.91) Seagem75.82 (2.93)82.43 (5.30)76.82 (2.24)85.30 (3.44)70.72 (6.51)63.64 (8.25)75.45 (1.47)81.39 (0.32) Table 8. F1 scores of fine-tuned models on SEAHateCheck and SGHateCheck High-Quality Test Cases, with changes from non- finetuned results in parentheses. Red = decrease, Blue = increase. decline on Mandarin, and others show drops on Malay or, in the case of Seagem, on both Indonesian and Mandarin. These variations suggest that finetuning does not uniformly stabilize multilingual performance and may reduce capabilities when the adaptation data is narrow or misaligned with the linguistic phenomena emphasized in SEAHateCheck and SGHateCheck. For example, SGHateCheck includes obfuscation types in Mandarin and Singlish—such as pinyin spellings, character decomposition, and spacing variants—that may be underrepresented in finetuning corpora, leading to performance deterioration. Comparing the two tables highlights that finetuned open models reduce the gap with closed baselines and, in some cases, surpass them on specific languages, as observed for Vietnamese and Singlish with Gemma. In addition, the most significant improvements occur in languages with the weakest out-of-the-box coverage, particularly Tagalog and Thai, reflecting the value of additional supervised signal in low-resource settings. From a model-centric perspective, architecture and pretraining scope can help to explain several observed effects. SeaLion and Seagem benefit from continual pretraining on regional data, which likely contributes to their strong out- of-the-box performance in Indonesian, Thai, and Vietnamese. Finetuning further enhances performance by providing targeted domain exposure. As shown in Table 16, Pangea excludes Tagalog in its base multilingual pretraining, accounting for its weak Tagalog baseline and significant improvements after adaptation. Larger instruction-tuned models such as Gemma and Qwen demonstrate broad gains, consistent with cross-lingual lexicalization and instruction-following capabilities acquired during pretraining. Conversely, performance declines on Mandarin for SeaLion, Seagem, Pangea, and Qwen suggest that finetuning can induce shifts in decision boundaries or cause forgetting when adaptation data is not aligned with the obfuscation patterns present in the benchmark. These effects reflect the challenge of modeling culturally grounded and orthographically diverse phenomena beyond simple word matching. From a language-centric perspective, the results align closely with the benchmarks’ design. SEAHateCheck targets implicit hate, slang, and culturally specific cues in Indonesian, Tagalog, Thai, and Vietnamese. At the same time, SGHateCheck extends coverage to Singapore-specific varieties, including Mandarin and Singlish, with additional obfuscation tests. Lower out-of-the-box performance for Tagalog and Tamil reflects both limited pretraining exposure and the intrinsic difficulty of counter-speech and negation contrasts. The substantial gains after finetuning indicate that modest, well-curated supervision can meaningfully enhance functional competence. In contrast, declines in Mandarin and specific Malay settings highlight domain mismatches between adaptation corpora and obfuscation-heavy test cases, emphasizing the need for explicit coverage of these phenomena to maintain robustness. Manuscript submitted to ACM 12Ng et al. Overall, the results demonstrate that finetuning is an effective strategy for SEA and Singaporean language varieties; however, its benefits depend on alignment with the functional phenomena being evaluated. Regressions can often be attributed to label imbalance, distributional skew in adaptation sets, forgetting of multilingual and orthographic knowledge, and insufficient representation of counter-speech or obfuscation examples central to SEAHateCheck and SGHateCheck. Future adaptation efforts should balance hateful and non-hateful supervision across languages, incorporate curated counter-speech and obfuscation examples, and employ strategies such as rehearsal or multi-task learning to preserve pretraining knowledge while specializing to culturally grounded hate expressions. 4.2 Performance across Functional Tests We summarize performance across functional tests below using a representative plot Figure 1, while the complete per- language radar plots are provided in Appendix Figures 7 - Figures 9. The non-finetuned models display a consistent profile: they are confident on explicit hate expressions yet brittle whenever the task requires discourse-level interpretation, reference tracking, or polarity reversal. On the explicit side, tests such as direct threat, slur usage, and profanity as hate remain near the top for most systems and languages, producing tightly clustered traces in the upper band of the radar plots. In contrast, non-hateful counter-speech is routinely mishandled. Performance is lowest on F18 and F19, which require recognizing quotations or references to hateful content to condemn it. This weakness is theoretically expected because the surface form of a counter-speech statement often reuses hateful tokens, inviting a spurious lexical shortcut. Our figures replicate this failure mode, as reported in the framework specification and narrative analysis, which already noted that models conflate refutation with hate and that F18 and F19 are especially challenging across languages. Negation also remains a persistent frontier. Models underperform on F13, where a hateful predicate is explicitly negated. The task requires aligning scope and compositional semantics instead of treating the presence of a hateful lexeme as decisive evidence, a pattern again anticipated in the task design and highlighted in the prior discussion of functional-level errors. Beyond counter-speech and negation, we observe variability in reference-based phenomena. Reference in subsequent clauses (F10) pulls many non-finetuned traces downward, and implicit derogation (F4) is uneven, especially in languages where pragmatic implicature relies on culture-specific idioms. These observations align with the framework’s emphasis on contextual reasoning beyond surface markers. Comprehensive per-function breakdowns for non-finetuned and finetuned models appear in Appendices D and E, respectively. Fine-tuning narrows several gaps yet does not eliminate them. The clearest gains appear on Indonesian, Thai, Vietnamese, Malay, and Singlish for the two counter-speech tests, where scores rise but often remain below chance for multiple model families, indicating partial rather than full transfer of discourse-level cues. This pattern matches the text of our study that reports post-adaptation improvements on F18 and F19 while cautioning that the absolute accuracy still lags behind explicit categories. In several settings, fine-tuning produces regressions on implicit and referential hate, notably F4, F10, and F12 for Indonesian and additional drops for Thai, Malay, and Mandarin. The combined evidence suggests that adaptation sets are rich in overt hate and categorical mentions but relatively sparse in examples that require long-range reference resolution or careful handling of polarity. When coupled with limited counter-speech coverage, this skew encourages over-reliance on lexical heuristics, a mechanism already discussed in our dataset-level analysis. Obfuscation tests illustrate another divide. Character-level perturbations in F23 to F27 and F32 to F33 depress several traces in the non-finetuned condition. Finetuning helps when the adaptation data expose the model to similar orthographic noise; however, improvements are inconsistent across models and languages, mirroring the framework’s design, which stresses robustness to creative evasion strategies found in real discourse. In contrast, non-protected-target Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia13 abuse (F20 to F22) sits higher and more stably both before and after fine-tuning, suggesting that models can usually separate personal or object-directed profanity from protected-class hate when group semantics are not at stake. Language-specific traces echo these global tendencies while revealing culturally grounded nuances. The Thai and Vietnamese plots exhibit broad post-finetuning gains around explicit threats and profanity, but smaller improvements on counter-speech and negation. Indonesian exhibits sizable improvements on reference and quotation counter-speech after finetuning, though implicit derogation remains fragile. Mandarin and Malay display sharper drops on some implicit and reference tests, consistent with concerns about domain mismatch between adaptation corpora and our obfuscation-heavy and discourse-sensitive test suites. These observations are consistent with the paper’s textual summary of per-language effects and with the caution that the benefits of finetuning depend on the alignment between adaptation phenomena and evaluation functions. From an experimental perspective, the radar plots reveal that fine-tuning narrows the spread of results across functional categories, indicating more stable behavior across models. Yet, differences persist across language groups. Vietnamese and Indonesian exhibit stronger results overall, reflecting the relative availability of training resources and clearer lexical markers. In contrast, Tagalog and Thai show wider variance, suggesting that linguistic complexity and cultural particularities make these cases more complicated for models to resolve. These cross-linguistic discrepancies underscore the necessity of SEAHateCheck by systematically diagnosing weaknesses at the functional test level. The dataset identifies precisely where models fail, whether in handling implicit hate, disambiguating non-hateful contrasts, or managing counter-speech. Taken together, the functional analysis indicates that current multilingual and regional models have largely solved explicit lexical hate under controlled conditions, while reasoning-heavy contrasts remain open challenges. Finetuning is effective when the adaptation data emphasize counter-speech, negation, and obfuscation. Still, it may erode performance on implicit or referential hate if the additional supervision overweights surface cues and induces forgetting of multi- lingual or orthographic knowledge. These conclusions reinforce SEAHateCheck’s value. By decomposing evaluation into complementary functions that reflect real moderation needs, it reveals where progress stems from vocabulary memorization and where genuine language understanding is still required. Manuscript submitted to ACM 14Ng et al. Fig. 1. Accuracy across Functional Tests for Indonesian (left) and Tagalog (right) 4.3 Performance across Protected Categories We evaluate performance across protected categories with a representative radar plot (Figure 2), and provide the full set of per-language results in Appendix Figures Figure 13 - Figure 15. The radar plots comparing non-finetuned and finetuned models consistently show that, while most systems achieve relatively high recall for categories such as Religion and Ethnicity/Race/Origin, their performance deteriorates substantially when handling more nuanced or culturally sensitive categories, including Gender/Sexuality, Disability, Age, People Living with HIV (PLHIV), and Vulnerable Workers. This trend is particularly salient in low-resource languages where linguistic complexity and sociocultural nuance exacerbate model weaknesses. Expanded category-wise results are available in Appendices F and G for non-finetuned and finetuned settings. Non-fine-tuned models tend to overfit towards high-frequency protected categories. For instance, across all tested languages, detection rates for Religion and Ethnicity/Race/Origin consistently cluster at the upper range of the radar charts, reflecting models’ reliance on explicit lexical cues and more frequent representation in training corpora. However, categories such as Gender/Sexuality and Disability show lower and more variable scores, highlighting that models Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia15 struggle with implicit expressions and culturally specific markers of marginalization. These findings align with prior research showing that functional tests involving implicit hate or counter-speech are among the most challenging. Fine-tuning improves overall balance across categories, particularly for Disability and Age, where finetuned models show more consistent recall across languages. Nonetheless, the gains are uneven. In categories like PLHIV and Vulnerable Workers, improvements remain limited, suggesting that fine-tuning with generic hate speech corpora does not sufficiently capture the cultural and legal salience of these groups in Southeast Asia. This phenomenon highlights the necessity of SEAHateCheck, which explicitly encodes these protected categories based on regional legislation and sociolinguistic consultation, thereby filling a critical gap absent in prior benchmarks such as HateCheck and Multilingual HateCheck. Larger or safety-hardened closed models (o3, Gemini) trace the outer envelope across categories before and after adaptation. Yet, even these systems show dips on Age and PLHIV in some languages, pointing to limited pretraining coverage of community-specific references and euphemisms. Open models benefit most from fine-tuning but remain sensitive to the distribution of training labels by category; when adaptation data underrepresents benign identity mentions, performance on Religion or Gender/Sexuality neutral cases can degrade due to over-triggering. Linguistically, tonal languages (Thai, Vietnamese) perform well in categories where explicit hate is common. Yet, they present additional challenges for Gender/Sexuality and Religion, where pragmatic markers and culture-bound idioms modulate toxicity. After finetuning, Thai and Vietnamese show clearer, more uniform polygons across categories, suggesting that localized templates capture tonal orthography that are rare in generic multilingual corpora. In contrast, Tagalog test cases reveal lower accuracy across nearly all categories, consistent with earlier findings that Tagalog’s code-switching and colloquial profanity are difficult for multilingual models to parse. However, it similarly tightens around Age and Disability once slangy references are included in training materials. These divergences highlight the fact that language-specific sociolinguistic phenomena directly impact category-level detection performance, rendering cross-lingual transfer insufficient without culturally grounded benchmarks. SEAHateCheck thus contributes to the field by systematically exposing model weaknesses across legally and culturally protected categories in Southeast Asia. It demonstrates that high overall accuracy on traditional test sets can mask severe blind spots in categories most relevant to marginalized communities. By providing contrastive and culturally validated functional tests, the dataset enables researchers to move beyond aggregate performance metrics and interrogate whether systems truly generalize across all vulnerable groups. In doing so, SEAHateCheck establishes itself as an indispensable resource for building inclusive hate speech detection systems and advancing equitable online content moderation in low-resource linguistic contexts. Manuscript submitted to ACM 16Ng et al. Fig. 2. F1 Score across Protected Categories on Indonesian (left) and Taglog (right) 5 Discussion on Silver Label Test Cases 5.1 Performance on Silver Testcases The silver test cases reveal performance dynamics that are not visible in gold cases, offering an essential diagnostic layer for assessing robustness. As shown in Tables 9 and 10, non-finetuned models perform competitively in some languages, such as Vietnamese (o3: 84.94, Gemini: 77.05, Minstral: 81.59) and Indonesian (o3: 81.52, Seagem: 76.32), where F1 scores exceed 75. In contrast, Tagalog and Tamil remain underexplored by most systems, with several models dropping below 60 (e.g., Deepseek: 53.86 on Tagalog, 56.37 on Tamil), underscoring the necessity of additional evaluation coverage through silver cases. Fine-tuning introduces striking gains in low-resource settings. Tagalog shows the most significant jumps, with Llama3b rising from 60.79 to 76.57 and Llama8b from 67.58 to 72.07, demonstrating the positive impact of even modest supervision in morphologically complex and underrepresented languages. Tamil similarly benefits, with Gemini climbing from 73.95 to 80.16, while Seagem improves from 74.78 to 77.10. These improvements validate the role of silver cases in surfacing progress where pretraining exposure is minimal and culturally grounded evaluation is scarce. At the Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia17 same time, regressions appear in better-resourced languages. Indonesian scores drop for most finetuned models, with Minstral declining from 76.27 to 66.34, SeaLion by 11.21, and Seagem by 13.44. Similar declines are observed in Thai, where Minstral drops by 4.67 and SeaLion by 6.34. These regressions suggest that adaptation corpora skewed toward explicit abuse may narrow models’ coverage and reduce calibration on colloquial or implicit cases that the silver sets capture. Chinese and Malay also show smaller but consistent decreases (e.g., Gemma shows−2.16 on Mandarin,−1.40 on Malay), which further highlights the sensitivity of multilingual systems to distributional mismatch. We provide comprehensive break-downs by functional test and protected category for both non-fine-tuned and fine-tuned models in Appendices D - G. From the model perspective, closed models sustain strong averages but their advantage is less pronounced on silver cases. For instance, o3 drops in Vietnamese from 84.94 to 78.12, showing that increased variability in case construction reduces the gap to open models. Conversely, Gemma and Qwen demonstrate broad adaptability, with Qwen improving in Tagalog and Tamil. At the same time, Gemma registers consistent, if modest, gains across several SEA languages (+3.41 in Indonesian, +4.05 in Thai, +4.90 in Vietnamese). This indicates that instruction-tuned models can leverage even limited supervision to extend generalisation to naturalistic, culturally embedded contexts. From the language perspective, the results track well with sociolinguistic characteristics. Vietnamese and Indonesian maintain high ceilings due to strong online representation and lexical regularity, although the observed declines after fine-tuning highlight the difficulty of balancing explicit and implicit phenomena. Tagalog and Tamil show the most apparent benefit from silver evaluation, since their underrepresentation in pretraining leaves models brittle to morphological richness and register variability. Singlish and Mandarin remain challenging: while fine-tuning yields higher recall, this comes with increased false positives on benign slang, obfuscations, or counter-speech, reflecting the cultural and orthographic ambiguity inherent to these varieties. Overall, the results demonstrate that silver test cases are indispensable in advancing hate speech detection for low-resource languages. They introduce variability, colloquialism, and length that expose weaknesses hidden by gold cases, while also providing a cost-efficient means to scale coverage. By complementing gold cases, the silver suite ensures that evaluation is not restricted to templated contrasts but extends into the messy, lived realities of online discourse in Southeast Asia. This contribution strengthens SEAHateCheck’s role as a benchmark. Model SEASG Indonesian Tagalog Thai Vietnamese Malay Mandarin Singlish Tamil Ministral76.2770.4266.8581.5968.9768.9766.3356.93 Llama3b62.4660.7963.0569.7061.5861.5860.9568.98 Llama8b69.7767.5868.4379.4469.0669.0668.7062.22 Sealion68.3764.5468.8778.6667.7767.7774.1564.26 Seallm58.2862.2066.0570.9465.4265.4267.0856.75 Pangea67.6860.8648.4269.4365.2665.2656.3751.84 Qwen75.9170.5873.6480.0271.9271.9269.0571.73 Gemma63.5261.3970.3978.5466.5066.5070.9356.21 Seagem76.3271.9773.4980.0072.7872.7870.4874.78 Gemini75.2964.0671.8377.0575.9875.9879.5676.30 o381.5274.8077.1784.9478.1278.1279.3879.98 Deepseek59.8353.8658.4966.6864.0164.0165.6062.37 Table 9. F1 scores of different non-finetuned models on SEAHateCheck and SGHateCheck Silver Test Cases. Manuscript submitted to ACM 18Ng et al. Model SEASG Indonesian TagalogThaiVietnameseMalayMandarinSinglishTamil Ministral66.34 (9.93)74.34 (3.92)71.52 (4.67)85.37 (3.78)70.65 (1.68)69.97 (1.00)75.68 (9.35)62.56 (5.63) Llama3b70.17 (7.71)76.57 (15.78)74.38 (11.33)85.98 (16.28)74.92 (13.34)71.82 (10.24)80.16 (19.21)75.19 (6.21) Llama8b57.83 (11.94)72.07 (4.49)61.43 (7.00)83.69 (4.25)68.05 (1.01)66.90 (2.16)73.51 (4.81)68.19 (5.97) Sealion57.16 (11.21)72.25 (7.71)65.46 (3.41)83.19 (4.53)63.63 (4.14)66.18 (1.59)70.87 (3.28)65.63 (1.37) Seallm61.09 (2.81)73.94 (11.74)73.99 (7.94)84.88 (13.94)73.94 (8.52)71.67 (6.25)79.69 (12.61)63.26 (6.51) Pangea61.37 (6.31)72.42 (11.56)63.21 (14.79)84.12 (14.69)69.91 (4.65)67.08 (1.82)72.48 (16.11)63.61 (11.77) Qwen69.64 (6.27)75.67 (5.09)71.14 (2.50)86.32 (6.30)73.32 (1.40)69.50 (2.42)81.00 (11.95)78.50 (6.77) Gemma66.93 (3.41)72.33 (10.94)69.34 (1.05)83.44 (4.90)75.40 (8.90)66.15 (0.35)76.46 (5.53)64.06 (7.85) Seagem62.88 (13.44)74.83 (2.86)68.14 (5.35)86.34 (6.34)69.68 (3.10)65.01 (7.77)79.49 (9.01)77.10 (2.32) Table 10. F1 scores of different fine-tuned models on SEAHateCheck and SGHateCheck Silver Test Cases, with changes from non-finetuned results in parentheses. Red = decrease, Blue = increase. To assess generalization on silver functional tests, we show a representative plot in Figure 3, Appendix Figure 12 - Figure 10 contain the remaining per-language results. Unlike gold test cases, which are template-driven and highly controlled, silver test cases are LLM-generated to reflect colloquial usage and richer linguistic diversity, making them a stronger proxy for real-world hate speech in Southeast Asian contexts. Across non-finetuned models, performance was markedly inconsistent. While several models demonstrated rea- sonable accuracy on explicit hate categories—such as derogatory negative emotion and attributional insults—their performance deteriorated on more subtle or implicit forms of hate speech. In particular, functional tests involving implicit derogation (F4), negated positive statements (F12), and phrasings in the form of questions or opinions (F14–F15) yielded much lower scores. These weaknesses reflect a reliance on surface-level lexical cues, which are less reliable in detecting implicit hate that is highly dependent on cultural and contextual interpretation. Similarly, non-finetuned models frequently misclassified counter-speech instances (F18–F19), where hate is quoted or explicitly denounced, highlighting persistent confusion between hateful and non-hateful expressions that share similar lexical markers. Fine-tuned models showed substantial improvements across nearly all functional categories, confirming the value of domain-specific adaptation. The most notable gains appeared in implicit hate detection and counter-speech recognition. For instance, models after fine-tuning achieved consistently higher F1-scores in the range of 0.7–0.9 for implicit derogation and counter-speech, whereas non-finetuned models often fell below 0.5. This indicates that fine-tuning improved models’ sensitivity to contextual cues and reduced false positives. Fine-tuned models also performed better on profanity-based tests (F8–F9), successfully distinguishing hateful versus non-hateful profanity—a contrast that non-finetuned models often failed to capture. Nevertheless, even fine-tuned models struggled with functional tests designed to capture identity references across subsequent sentences or clauses (F10–F11), as well as complex substitution-based tests (F23–F34). These involve longer discourse structures or obfuscation patterns that require robust contextual reasoning and adaptability. The gap between performance on template-based gold cases and silver cases also illustrates the difficulty of transferring model competence to naturally occurring, diverse inputs. The silver functional tests underscore the necessity of SEAHateCheck for low-resource Southeast Asian languages. They expose failure modes that are unlikely to surface in template-driven evaluation, particularly in handling collo- quial phrasing, cultural nuances, and implicit hate that dominate real-world discourse. The diagnostic results high- light how fine-tuning mitigates specific weaknesses, yet also reveal areas where further research is needed—such as Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia19 discourse-aware architectures and cross-sentence contextual modeling. By incorporating both gold and silver test cases, SEAHateCheck enables a more comprehensive evaluation of hate speech detection systems, bridging the gap between controlled diagnostic testing and the variability of authentic online environments. 5.2 Performance across Protected Categories on Silver Testcases We examine silver-case performance across protected categories using a representative plot (Figure 4), with additional per-language plots in Appendix Figure 18 - Figure 16. Overall, both non-finetuned and finetuned models exhibited consistent trends across categories such as Religion, Ethnicity/Race/Origin, Gender/Sexuality, Disability, Age, Vulnerable Workers, and People Living with HIV (PLHIV). These categories were defined based on local legislation (see Section 2.1), making them critical for assessing whether detection models are culturally attuned to the sociopolitical realities of Southeast Asia. From a qualitative perspective, non-finetuned models tended to perform unevenly across categories. Religion and Ethnicity/Race/Origin generally achieved higher detection rates, likely because these categories are more frequently represented in global pretraining corpora. In contrast, categories such as Disability, PLHIV, and Vulnerable Workers proved particularly challenging, with models often failing to recognize implicit or culturally nuanced hateful expressions. This difficulty reflects both the scarcity of such examples in training data and the sociolinguistic complexity of localized insults and stereotypes. Fine-tuning improved robustness, reducing variance across categories. However, even after fine-tuning, specific categories such as Disability and PLHIV remained relatively underperforming compared to Religion or Ethnicity, suggesting that fine-tuning alone cannot fully mitigate the gaps caused by limited representation in upstream datasets. Quantitatively, the silver testcases highlight substantial differences in accuracy across protected groups. In non- finetuned settings, model performance on Religion and Ethnicity averaged between 0.70 and 0.78 F1, while scores for Disability and PLHIV often fell below 0.60. Fine-tuned models exhibited an overall upward shift, with Religion and Ethnicity surpassing 0.80 F1 in most cases and Gender/Sexuality stabilizing around 0.75. Nevertheless, gains were uneven: while Disability improved from 0.55 to 0.68 and PLHIV from 0.52 to 0.65, these categories still lagged behind better-represented groups. Age and Vulnerable Workers presented mixed outcomes, with some models (e.g., SEA-Lion, Seallm) achieving performance comparable to Ethnicity (0.72–0.74), whereas others plateaued closer to 0.65. This pattern underscores the persistent data imbalance across categories. Taken together, these findings reinforce the necessity of SEAHateCheck’s protected-category-specific functional tests. By exposing systematic weaknesses across underrepresented categories such as Disability, PLHIV, and Vulnerable Workers, SEAHateCheck provides diagnostic insights that are often obscured by aggregate accuracy metrics. The silver testcases, in particular, amplify this diagnostic power by introducing more naturalistic and colloquial expressions that stress-test models beyond template-based gold cases. Beyond label accuracy, recent work proposes metrics for evaluating the reasoning quality of hate-speech explanations, offering an additional diagnostic lens for trustworthy moderation [20]. Such metrics could complement SEAHateCheck-style functional testing in future evaluations. Manuscript submitted to ACM 20Ng et al. Fig. 3. Accuracy across Silver Functional Tests for Indonesian (left) and Tagalog (right) Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia21 Fig. 4. F1 Score across Protected Categories for Silver Indonesian (left) and Taglog (right) 6 Related Work Research on online hate speech detection has been propelled by annotated resources that concentrate on English and a handful of other well-resourced languages, which accelerates model progress but obscures weaknesses when models face culturally distinct phenomena in low-resource settings. This imbalance is acute in Southeast Asia, where heterogeneous scripts, tone-sensitive phonology, code mixing, and country-specific legal framings of protected groups interact with platform discourse in ways that are not captured by existing benchmarks. A dataset for this region must therefore provide culturally grounded evaluation with interpretable, capability-level diagnostics while remaining methodologically transparent and reproducible. Functional testing offers a principled alternative to random held-out evaluation by probing well-defined capabilities such as handling negation, distinguishing benign identity mentions from abuse, and recognizing counterspeech. HateCheck introduced this paradigm in English through hand crafted functional tests that expose systematic errors that aggregate accuracy can hide [41]. Multilingual HateCheck extended the idea to ten languages and showed that even Manuscript submitted to ACM 22Ng et al. strong multilingual systems fail in predictable ways when confronted with targeted linguistic phenomena Röttger et al. [40]. These works motivate a shift toward interpretable evaluation suites that can reveal failure modes tied to culture and language rather than only corpus level statistics. In the Southeast Asian context, Hu et al. [19]complements our functional-test perspective with adversarial safety evaluation. This further motivates region-specific resources beyond English-centric benchmarks. A common strategy for expanding coverage to new languages is translation. Some resources rely on human translation to preserve nuance, as in Röttger et al. [40]. Others adopt machine translation to scale the process; for example, Goldzycher et al. [17]uses Google Translate 2 to create candidate instances that are subsequently curated. SGHateCheck follows a hybrid approach by using large language models to assist translation and paraphrasing while incorporating feedback from two native speakers to ensure fidelity [30]. Our work builds upon this pipeline and advances it by using few-shot prompting seeded with previously verified translations, ensuring that machine-assisted outputs better reflect country-specific usage, taboo lexicon, and local discourse markers before being adjudicated by native experts. Another line generates entirely novel test cases with large language models. Shen et al. [45]constructs a broad set of hateful and non hateful content by combining toxic personas [13] with jailbreak style prompts [44], and reports a non trivial safety rejection rate as models decline to produce some requested content. By contrast, the silver label test cases in SEAHateCheck and in SGHateCheck are produced with prompts co designed with native speakers and conditioned on high quality gold examples from both suites. We deliberately avoid jailbreak tactics and instead employ SEA Lionv2.1 due to its relatively low rejection rate under safety guardrails, which yields colloquial yet controlled variants aligned with explicit functions rather than unconstrained toxicity. Recent work asks whether large language models can author functional tests themselves. GPT HateCheck studies prompt and rubric designs for generating capability targeted minimal pairs and reports that LLM authored tests can diversify phrasing and expand coverage when strong quality control is imposed, while also noting that models tend to introduce culturally brittle phrasing and subtle shifts in intent when left unchecked Jin et al. [23]. SEAHateCheck adopts a conservative position in light of these findings. Gold cases remain human curated and template instantiated to guarantee interpretability and legal alignment, and silver cases use LLM generation only under native co-design with few-shot seeding, followed by unanimous adjudication. This choice retains the diversity benefits highlighted by GPT HateCheck and mitigates risks from model priors that are poorly calibrated for Southeast Asian discourse. Large language models are increasingly used as detectors rather than only as generators. Pan et al. [36]evaluates instruction following models with zero-shot and few-shot prompting and with parameter-efficient fine-tuning via LoRA [18], observing that fine-tuned variants can gain recall at the expense of precision. In a complementary study, Das et al. [7]shows that a simple prompt to a state of the art conversational model performs competitively on aggregate yet struggles with counterspeech and with non English inputs. SEAHateCheck directly targets these weaknesses by including contrasts for counterspeech, quotation, and denouncement, benign identity mentions, and profanity as a discourse marker, thereby offering fine-grained diagnostics for both prompted and fine-tuned detectors. SEAHateCheck is strongly inspired by SGHateCheck but addresses limitations that arise from a Singapore-centric scope [30]. Both suites embrace capability-focused testing with expert-in-the-loop validation, and both pair template- instantiated gold items with silver items that widen lexical and compositional variety. SEAHateCheck expands language and country coverage to Indonesia, the Philippines, Thailand, and Vietnam, introduces tone-sensitive languages and distinct writing systems, and regrounds protected categories through consultation with local experts and legislation. 2 https://translate.google.com/ Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia23 The multi-stage pipeline combines curation of machine-assisted translations with unanimous adjudication to ensure unambiguous intent, while few-shot silver generation seeded by verified gold exemplars captures colloquial forms, obfuscation strategies, and reclaimed or figurative expressions that frequently confuse detectors. This design preserves interpretability at the function level and increases the ecological validity of tests for the region. Beyond diagnostics, the dataset enables rigorous assessment of adaptation strategies for low-resource settings. Awal et al. [2]shows that model agnostic meta learning yields initializations that adapt quickly to unseen languages under limited supervision. Ghorbanpour et al. [16]demonstrates that cross-lingual nearest neighbor retrieval can leverage small labeled sets by augmenting them with retrieved neighbors from multilingual pools. Ye et al. [56]explores privacy- preserving few-shot learning with datasets curated by marginalized communities, which is crucial where centralizing sensitive text is infeasible. Evaluating these methods on SEAHateCheck can disentangle actual improvements in capability from distribution-specific shortcuts by measuring gains on targeted functions such as implicit abuse, negation, reference, and slang. Community-centered scholarship further argues that local expertise is essential for reliable low-resource hate speech pipelines. Nkemelu et al. [31]shows that context experts are needed to define protected groups, taboo lexicons, and pragmatic cues that shape interpretation. SEAHateCheck operationalizes this guidance by embedding native translators and adjudicators throughout template adaptation, by documenting country-specific categories and slur inventories, and by instituting qualitative debriefs to capture cultural misfires for subsequent revision, which together reduce conflation of routine slang with genuine abuse. Methodological breadth continues to expand with representation learning and semi-supervised strategies that our dataset can stress test. A dual contrastive framework is introduced to improve discrimination under data scarcity by aligning representations of hateful and non-hateful content [6]. Transformer-based systems with explainable components provide transparency alongside performance gains [15], while semi-supervised generative adversarial approaches leverage unlabeled data to improve generalization across languages [29]. Beyond neural and multilingual approaches several studies analyse data pre-processing pipelines and feature engineering strategies for hate speech detection including work on women targeted abuse and feature combinations on Twitter [39,42]. Because SEAHateCheck supplies carefully matched minimal pairs and explicit function labels, it enables precise evaluation of whether contrastive, explainable, or generative systems capture intent rather than correlating on surface profanity or identity tokens. Finally, the field has begun to consolidate insights across languages and methods. Recent benchmarks examine culturally grounded evaluation across Asian contexts, indicating a broader trend toward culture-aware evaluation beyond Western high-resource settings [59]. Das et al. [8]reviews resources and techniques for low-resource hate speech detection and highlights two persistent gaps, namely the scarcity of culturally grounded datasets that reflect code mixing and proverb-based insinuation, and the lack of evaluation protocols that diagnose capabilities rather than only reporting corpus-level scores. By delivering a functionally controlled and culturally anchored suite for several Southeast Asian languages, SEAHateCheck addresses both deficits and contributes a common yardstick for future models that span meta learning, retrieval, federated optimization, contrastive objectives, and semi-supervised generation. In this way the dataset advances the empirical foundation for reliable detection in a linguistically and culturally diverse region. 7 Conclusion This study introduces SEAHateCheck, a HS benchmark dataset comprising testcases curated to the sociocultural landscape of Indonesia, the Philippines, Thailand, and Vietnam, and is designed to test HS detectors on a variety of HS functionalities and target groups. SEAHateCheck is further split into Gold Label and Silver Label datasets. Gold Manuscript submitted to ACM 24Ng et al. Label datasets are translated from HateCheck [41], and three annotators for each testcase annotated selected ones. Silver Label datasets were generated by few-shot learning using high-quality test cases identified by annotators. This procedure is also used to form the Silver Label dataset for SGHateCheck [30]. In general, we observe that Silver Label testcases are more extensive and less likely to be judged by annotators as high quality. Subsequently, Gold Label and Silver Label SEAHateCheck and SGHateCheck testcases were tested on LLMs, where they were prompted to behave like HS detectors. In general, both closed and open-sourced LLMs performed exceptionally well and frequently had an F1 of above 0.7. That said, we observed varying degrees of performance across different languages, functional tests, and protected categories. Fine-tuning also led to higher precision for most cases, but also increased the risk of overtraining. 8 Limitations A major limitation of SEAHateCheck Gold Label dataset is the rigidity of using templates to generate testcases, which limits the ability to customise templates to specific targets. The Silver Label datasets, where HS was machine generated with input from the Gold Label dataset, were designed to overcome this rigidity issue. However, native-speakers were more likely to find the Silver Label datasets to be of lower quality. A possible solution would be to use more powerful LLMs with jailbroken prompts, as described in Shen et al. [45]. Where silver quality scores fall short of gold and where target or function alignment is imperfect, we provide language-specific adjudication notes in Appendix B that motivated subsequent filtering and template revisions. Recent work shows that cloaking perturbations can substantially degrade offensive-language detection robustness, suggesting that evasion-focused transformations should be incorporated into future SEAHateCheck extensions [54]. The structure of the templates, which have up to two sentences, also does not fully reflect the conversational nature of interactions on social media. As our experiments have shown, recent advancement in LLMs has made HS detection for such short texts a trivial matter. Text that seem harmless alone, when chained together, can result in toxic meanings [48]. Another major limitation of this study is the reliance on existing laws to identify target groups, which represents a lag in vulnerable groups in society today. Acknowledgments This research/project is supported by Ministry of Education, Singapore, under its Academic Research Fund (AcRF) Tier 2. Any opinions, findings and conclusions or recommendations expressed in this material are those of the authors and do not reflect the views of the Ministry of Education, Singapore. References [1] 2024. Un Ministral, des Ministraux. https://mistral.ai/news/ministraux. [2] Md Rabiul Awal, Roy Ka-Wei Lee, Eshaan Tanwar, Tanmay Garg, and Tanmoy Chakraborty. 2023. Model-agnostic meta-learning for multilingual hate speech detection. IEEE Transactions on Computational Social Systems 11, 1 (2023), 1086–1095. [3]Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 1877–1901. https://proceedings.neurips.c/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf [4] Neil Vicente Cabasag, Vicente Raphael Chan, Sean Christian Lim, Mark Edward Gonzales, and Charibeth Cheng. 2019. Hate speech in Philippine election-related tweets: Automatic detection and classification using natural language processing. Philippine Computing Journal XIV, 1 (August 2019). Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia25 [5]Bharathi Raja Chakravarthi, Vigneshwaran Muralidaran, Ruba Priyadharshini, and John Philip McCrae. 2020. Corpus Creation for Sentiment Analysis in Code-Mixed Tamil-English Text. In Proceedings of the 1st Joint Workshop on Spoken Language Technologies for Under-resourced languages (SLTU) and Collaboration and Computing for Under-Resourced Languages (CCURL), Dorothee Beermann, Laurent Besacier, Sakriani Sakti, and Claudia Soria (Eds.). European Language Resources association, Marseille, France, 202–210. https://aclanthology.org/2020.sltu-1.28/ [6]Krishan Chavinda and Uthayasanker Thayasivam. 2025. A Dual Contrastive Learning Framework for Enhanced Hate Speech Detection in Low- Resource Languages. In Proceedings of the First Workshop on Challenges in Processing South Asian Languages (CHiPSAL 2025), Kengatharaiyer Sarveswaran, Ashwini Vaidya, Bal Krishna Bal, Sana Shams, and Surendrabikram Thapa (Eds.). International Committee on Computational Linguistics, Abu Dhabi, UAE, 115–123. https://aclanthology.org/2025.chipsal-1.11/ [7] Mithun Das, Saurabh Kumar Pandey, and Animesh Mukherjee. 2024. Evaluating ChatGPT against Functionality Tests for Hate Speech Detection. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue (Eds.). ELRA and ICCL, Torino, Italia, 6370–6380. https://aclanthology.org/2024.lrec-main.564/ [8]Susmita Das, Arpita Dutta, Kingshuk Roy, Abir Mondal, and Arnab Mukhopadhyay. 2024. A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages. arXiv preprint arXiv:2411.19017 (2024). [9]Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated Hate Speech Detection and the Problem of Offensive Language. Proceedings of the International AAAI Conference on Web and Social Media 11, 1 (May 2017), 512–515. https://doi.org/10.1609/icwsm.v11i1.14955 [10]Google Deepmind. 2025. Introducing Gemini 2.0: our new AI model for the agentic era. https://blog.google/technology/google-deepmind/google- gemini-ai-update-december-2024/#gemini-2-0-flash [11]Fabio Del Vigna12, Andrea Cimino23, Felice Dell’Orletta, Marinella Petrocchi, and Maurizio Tesconi. 2017. Hate me, hate me not: Hate speech detection on facebook. In Proceedings of the First Italian Conference on Cybersecurity (ITASEC17). 86–95. arXiv:http://ceur-ws.org/Vol-1816/paper- 09.pdf http://ceur-ws.org/Vol-1816/paper-09.pdf [12]Jiawen Deng, Jingyan Zhou, Hao Sun, Chujie Zheng, Fei Mi, Helen Meng, and Minlie Huang. 2022. COLD: A Benchmark for Chinese Offensive Language Detection. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 11580–11599. https://doi.org/10.18653/v1/2022.emnlp- main.796 [13] Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023. Toxicity in chatgpt: Analyzing persona- assigned language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 1236–1270. https://doi.org/10.18653/v1/2023.findings-emnlp.88 [14]Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Lauren Rantala-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Olivier Duchenne, Onur Çelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, Vladan Petrovic, Weiwei Chu, Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaoqing Ellen Tan, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aaron Grattafiori, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alex Vaughan, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Franco, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Manuscript submitted to ACM 26Ng et al. Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl Parker, Carly Burton, Catalina Mejia, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, Danny Wyatt, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Firat Ozgenel, Francesco Caggioni, Francisco Guzmán, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Govind Thattai, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Irina-Elena Veliche, Itai Gat, Jake Weissman, James Geboski, James Kohli, Japhet Asher, Jean-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Karthik Prasad, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kun Huang, Kunal Chawla, Kushal Lakhotia, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Maria Tsimpoukelli, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Michael L. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikolay Pavlovich Laptev, Ning Dong, Ning Zhang, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Rohan Maheswari, Russ Howes, Ruty Rinott, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satterfield, Sudarshan Govindaprasad, Sumit Gupta, Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Kohler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Kumar, Vishal Mangla, Vítor Albiero, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaofang Wang, Xiaojian Wu, Xiaolan Wang, Xide Xia, Xilun Wu, Xinbo Gao, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yuchen Hao, Yundi Qian, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, and Zhiwei Zhao. 2024. The Llama 3 Herd of Models. arXiv:2407.21783 [cs.AI] https://arxiv.org/abs/2407.21783 [15] Endrit Fetahi, Arsim Susuri, Mentor Hamiti, Zenun Kastrati, Ercan Canhasi, and Arta Misini. 2025. Enhancing social media hate speech detection in low-resource languages using transformers and explainable AI. Social Network Analysis and Mining 15, 1 (2025), 82. [16] Faeze Ghorbanpour, Daryna Dementieva, and Alexander Fraser. 2025. Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neigh- bor Retrieval with Limited Labeled Data. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (Eds.). Association for Computational Linguistics, Suzhou, China, 29662–29680. https://doi.org/10.18653/v1/2025.emnlp-main.1507 [17] Janis Goldzycher, Paul Röttger, and Gerold Schneider. 2024. Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Kevin Duh, Helena Gomez, and Steven Bethard (Eds.). Association for Computational Linguistics, Mexico City, Mexico, 4405–4424. https://doi.org/10.18653/v1/2024.naacl-long.248 [18]Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations. https://openreview.net/forum?id=nZeVKeeFYf9 [19]Yujia Hu, Ming Shan Hee, Preslav Nakov, and Roy Ka-Wei Lee. 2025. Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore’s Low- Resource Languages. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (Eds.). Association for Computational Linguistics, Suzhou, China, 12183–12201. https: //doi.org/10.18653/v1/2025.emnlp-main.612 [20]Yujia Hu and Roy Ka-Wei Lee. 2026. HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations. arXiv preprint arXiv:2601.13547 (2026). [21] Muhammad Huzaifah, Weihua Zheng, Nattapol Chanpaisit, and Kui Wu. 2024. Evaluating Code-Switching Translation with Large Language Models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue (Eds.). ELRA and ICCL, Torino, Italia, 6381–6394. https://aclanthology.org/2024.lrec-main.565/ [22]Muhammad Okky Ibrohim and Indra Budi. 2019. Multi-label Hate Speech and Abusive Language Detection in Indonesian Twitter. In Proceedings of the Third Workshop on Abusive Language Online, Sarah T. Roberts, Joel Tetreault, Vinodkumar Prabhakaran, and Zeerak Waseem (Eds.). Association for Computational Linguistics, Florence, Italy, 46–57. https://doi.org/10.18653/v1/W19-3506 Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia27 [23]Yiping Jin, Leo Wanner, and Alexander Shvets. 2024. GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue (Eds.). ELRA and ICCL, Torino, Italia, 7867–7885. https://aclanthology.org/2024.lrec-main.694/ [24] Klaus Krippendorff. 2018. Content analysis: An introduction to its methodology. Sage publications. 357 pages. [25]Son T. Luu, Kiet Van Nguyen, and Ngan Luu-Thuy Nguyen. 2021. A Large-Scale Dataset for Hate Speech Detection on Vietnamese Social Media Texts. Springer International Publishing, 415–426. https://doi.org/10.1007/978-3-030-79457-6_35 [26]Krishanu Maity, Shaubhik Bhattacharya, Sriparna Saha, and Manjeevan Seera. 2023. A Deep Learning Framework for the Detection of Malay Hate Speech. IEEE Access 11 (2023), 79542–79552. https://doi.org/10.1109/ACCESS.2023.3298808 [27]Krishanu Maity, A. S. Poornash, Shaubhik Bhattacharya, Salisa Phosit, Sawarod Kongsamlit, Sriparna Saha, and Kitsuchart Pasupa. 2024. HateThaiSent: Sentiment-Aided Hate Speech Detection in Thai Language. IEEE Transactions on Computational Social Systems 11, 5 (2024), 5714–5727. https: //doi.org/10.1109/TCSS.2024.3376958 [28]Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021. Hatexplain: A benchmark dataset for explainable hate speech detection. In Proceedings of the AAAI conference on artificial intelligence, Vol. 35. 14867–14875. [29]Khouloud Mnassri, Reza Farahbakhsh, and Noel Crespi. 2024. Multilingual hate speech detection: a semi-supervised generative adversarial approach. Entropy 26, 4 (2024), 344. [30]Ri Chi Ng, Nirmalendu Prakash, Ming Shan Hee, Kenny Tsu Wei Choo, and Roy Ka-wei Lee. 2024. SGHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Singapore. In Proceedings of the 8th Workshop on Online Abuse and Harms (WOAH 2024), Yi-Ling Chung, Zeerak Talat, Debora Nozza, Flor Miriam Plaza-del Arco, Paul Röttger, Aida Mostafazadeh Davani, and Agostina Calabrese (Eds.). Association for Computational Linguistics, Mexico City, Mexico, 312–327. https://doi.org/10.18653/v1/2024.woah-1.24 [31]Daniel Nkemelu, Harshil Shah, Michael Best, and Irfan Essa. 2022. Tackling hate speech in low-resource languages with context experts. In Proceedings of the 2022 International Conference on information and communication technologies and development. 1–11. [32]OpenAI, :, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander Mądry, Alex Baker-Whitcomb, Alex Beutel, Alex Borzunov, Alex Carney, Alex Chow, Alex Kirillov, Alex Nichol, Alex Paino, Alex Renzin, Alex Tachard Passos, Alexander Kirillov, Alexi Christakis, Alexis Conneau, Ali Kamali, Allan Jabri, Allison Moyer, Allison Tam, Amadou Crookes, Amin Tootoochian, Amin Tootoonchian, Ananya Kumar, Andrea Vallone, Andrej Karpathy, Andrew Braunstein, Andrew Cann, Andrew Codispoti, Andrew Galu, Andrew Kondrich, Andrew Tulloch, Andrey Mishchenko, Angela Baek, Angela Jiang, Antoine Pelisse, Antonia Woodford, Anuj Gosalia, Arka Dhar, Ashley Pantuliano, Avi Nayak, Avital Oliver, Barret Zoph, Behrooz Ghorbani, Ben Leimberger, Ben Rossen, Ben Sokolowsky, Ben Wang, Benjamin Zweig, Beth Hoover, Blake Samic, Bob McGrew, Bobby Spero, Bogo Giertler, Bowen Cheng, Brad Lightcap, Brandon Walkin, Brendan Quinn, Brian Guarraci, Brian Hsu, Bright Kellogg, Brydon Eastman, Camillo Lugaresi, Carroll Wainwright, Cary Bassin, Cary Hudson, Casey Chu, Chad Nelson, Chak Li, Chan Jun Shern, Channing Conger, Charlotte Barette, Chelsea Voss, Chen Ding, Cheng Lu, Chong Zhang, Chris Beaumont, Chris Hallacy, Chris Koch, Christian Gibson, Christina Kim, Christine Choi, Christine McLeavey, Christopher Hesse, Claudia Fischer, Clemens Winter, Coley Czarnecki, Colin Jarvis, Colin Wei, Constantin Koumouzelis, Dane Sherburn, Daniel Kappler, Daniel Levin, Daniel Levy, David Carr, David Farhi, David Mely, David Robinson, David Sasaki, Denny Jin, Dev Valladares, Dimitris Tsipras, Doug Li, Duc Phong Nguyen, Duncan Findlay, Edede Oiwoh, Edmund Wong, Ehsan Asdar, Elizabeth Proehl, Elizabeth Yang, Eric Antonow, Eric Kramer, Eric Peterson, Eric Sigler, Eric Wallace, Eugene Brevdo, Evan Mays, Farzad Khorasani, Felipe Petroski Such, Filippo Raso, Francis Zhang, Fred von Lohmann, Freddie Sulit, Gabriel Goh, Gene Oden, Geoff Salmon, Giulio Starace, Greg Brockman, Hadi Salman, Haiming Bao, Haitang Hu, Hannah Wong, Haoyu Wang, Heather Schmidt, Heather Whitney, Heewoo Jun, Hendrik Kirchner, Henrique Ponde de Oliveira Pinto, Hongyu Ren, Huiwen Chang, Hyung Won Chung, Ian Kivlichan, Ian O’Connell, Ian O’Connell, Ian Osband, Ian Silber, Ian Sohl, Ibrahim Okuyucu, Ikai Lan, Ilya Kostrikov, Ilya Sutskever, Ingmar Kanitscheider, Ishaan Gulrajani, Jacob Coxon, Jacob Menick, Jakub Pachocki, James Aung, James Betker, James Crooks, James Lennon, Jamie Kiros, Jan Leike, Jane Park, Jason Kwon, Jason Phang, Jason Teplitz, Jason Wei, Jason Wolfe, Jay Chen, Jeff Harris, Jenia Varavva, Jessica Gan Lee, Jessica Shieh, Ji Lin, Jiahui Yu, Jiayi Weng, Jie Tang, Jieqi Yu, Joanne Jang, Joaquin Quinonero Candela, Joe Beutler, Joe Landers, Joel Parish, Johannes Heidecke, John Schulman, Jonathan Lachman, Jonathan McKay, Jonathan Uesato, Jonathan Ward, Jong Wook Kim, Joost Huizinga, Jordan Sitkin, Jos Kraaijeveld, Josh Gross, Josh Kaplan, Josh Snyder, Joshua Achiam, Joy Jiao, Joyce Lee, Juntang Zhuang, Justyn Harriman, Kai Fricke, Kai Hayashi, Karan Singhal, Katy Shi, Kavin Karthik, Kayla Wood, Kendra Rimbach, Kenny Hsu, Kenny Nguyen, Keren Gu-Lemberg, Kevin Button, Kevin Liu, Kiel Howe, Krithika Muthukumar, Kyle Luther, Lama Ahmad, Larry Kai, Lauren Itow, Lauren Workman, Leher Pathak, Leo Chen, Li Jing, Lia Guy, Liam Fedus, Liang Zhou, Lien Mamitsuka, Lilian Weng, Lindsay McCallum, Lindsey Held, Long Ouyang, Louis Feuvrier, Lu Zhang, Lukas Kondraciuk, Lukasz Kaiser, Luke Hewitt, Luke Metz, Lyric Doshi, Mada Aflak, Maddie Simens, Madelaine Boyd, Madeleine Thompson, Marat Dukhan, Mark Chen, Mark Gray, Mark Hudnall, Marvin Zhang, Marwan Aljubeh, Mateusz Litwin, Matthew Zeng, Max Johnson, Maya Shetty, Mayank Gupta, Meghan Shah, Mehmet Yatbaz, Meng Jia Yang, Mengchao Zhong, Mia Glaese, Mianna Chen, Michael Janner, Michael Lampe, Michael Petrov, Michael Wu, Michele Wang, Michelle Fradin, Michelle Pokrass, Miguel Castro, Miguel Oom Temudo de Castro, Mikhail Pavlov, Miles Brundage, Miles Wang, Minal Khan, Mira Murati, Mo Bavarian, Molly Lin, Murat Yesildal, Nacho Soto, Natalia Gimelshein, Natalie Cone, Natalie Staudacher, Natalie Summers, Natan LaFontaine, Neil Chowdhury, Nick Ryder, Nick Stathas, Nick Turley, Nik Tezak, Niko Felix, Nithanth Kudige, Nitish Keskar, Noah Deutsch, Noel Bundick, Nora Puckett, Ofir Nachum, Ola Okelola, Oleg Boiko, Oleg Murk, Oliver Jaffe, Olivia Watkins, Olivier Godement, Owen Campbell-Moore, Patrick Chao, Paul McMillan, Pavel Belov, Peng Su, Peter Bak, Peter Bakkum, Peter Deng, Peter Dolan, Peter Hoeschele, Peter Welinder, Phil Tillet, Philip Pronin, Philippe Tillet, Prafulla Dhariwal, Qiming Yuan, Rachel Dias, Rachel Manuscript submitted to ACM 28Ng et al. Lim, Rahul Arora, Rajan Troll, Randall Lin, Rapha Gontijo Lopes, Raul Puri, Reah Miyara, Reimar Leike, Renaud Gaubert, Reza Zamani, Ricky Wang, Rob Donnelly, Rob Honsby, Rocky Smith, Rohan Sahai, Rohit Ramchandani, Romain Huet, Rory Carmichael, Rowan Zellers, Roy Chen, Ruby Chen, Ruslan Nigmatullin, Ryan Cheu, Saachi Jain, Sam Altman, Sam Schoenholz, Sam Toizer, Samuel Miserendino, Sandhini Agarwal, Sara Culver, Scott Ethersmith, Scott Gray, Sean Grove, Sean Metzger, Shamez Hermani, Shantanu Jain, Shengjia Zhao, Sherwin Wu, Shino Jomoto, Shirong Wu, Shuaiqi, Xia, Sonia Phene, Spencer Papay, Srinivas Narayanan, Steve Coffey, Steve Lee, Stewart Hall, Suchir Balaji, Tal Broda, Tal Stramer, Tao Xu, Tarun Gogineni, Taya Christianson, Ted Sanders, Tejal Patwardhan, Thomas Cunninghman, Thomas Degry, Thomas Dimson, Thomas Raoux, Thomas Shadwell, Tianhao Zheng, Todd Underwood, Todor Markov, Toki Sherbakov, Tom Rubin, Tom Stasi, Tomer Kaftan, Tristan Heywood, Troy Peterson, Tyce Walters, Tyna Eloundou, Valerie Qi, Veit Moeller, Vinnie Monaco, Vishal Kuo, Vlad Fomenko, Wayne Chang, Weiyi Zheng, Wenda Zhou, Wesam Manassra, Will Sheu, Wojciech Zaremba, Yash Patil, Yilei Qian, Yongjik Kim, Youlong Cheng, Yu Zhang, Yuchen He, Yuchen Zhang, Yujia Jin, Yunxing Dai, and Yury Malkov. 2024. GPT-4o System Card. arXiv:2410.21276 [cs.CL] https://arxiv.org/abs/2410.21276 [33] OpenAI. 2025. OpenAI o3-mini. https://openai.com/index/openai-o3-mini/ [34] OpenAI. 2025. OpenAI o3-mini. https://openai.com/index/openai-o3-mini/ [35]Nedjma Ousidhoum, Zizheng Lin, Hongming Zhang, Yangqiu Song, and Dit-Yan Yeung. 2019. Multilingual and Multi-Aspect Hate Speech Analysis. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.). Association for Computational Linguistics, Hong Kong, China, 4675–4684. https://doi.org/10.18653/v1/D19-1474 [36]Ronghao Pan, José Antonio García-Díaz, and Rafael Valencia-García. 2024. Comparing Fine-Tuning, Zero and Few-Shot Strategies with Large Language Models in Hate Speech Detection in English. CMES-Computer Modeling in Engineering & Sciences 140, 3 (2024). [37]Fabio Poletto, Valerio Basile, Manuela Sanguinetti, Cristina Bosco, and Viviana Patti. 2021. Resources and benchmark corpora for hate speech detection: a systematic review. Lang. Resour. Eval. 55, 2 (June 2021), 477–523. https://doi.org/10.1007/s10579-020-09502-8 [38]Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu. 2025. Qwen2.5 Technical Report. arXiv:2412.15115 [cs.CL] https://arxiv.org/abs/2412.15115 [39] Rutuja G. Rathod, Yashoda Barve, Jatinderkumar R. Saini, and Sourav Rathod. 2023. From Data Pre-processing to Hate Speech Detection: An Interdisciplinary Study on Women-targeted Online Abuse. In 2023 3rd International Conference on Intelligent Technologies (CONIT). 1–8. https://doi.org/10.1109/CONIT59222.2023.10205571 [40]Paul Röttger, Haitham Seelawi, Debora Nozza, Zeerak Talat, and Bertie Vidgen. 2022. Multilingual HateCheck: Functional Tests for Multilingual Hate Speech Detection Models. In Proceedings of the Sixth Workshop on Online Abuse and Harms (WOAH), Kanika Narang, Aida Mostafazadeh Davani, Lambert Mathias, Bertie Vidgen, and Zeerak Talat (Eds.). Association for Computational Linguistics, Seattle, Washington (Hybrid), 154–169. https://doi.org/10.18653/v1/2022.woah-1.15 [41] Paul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert. 2021. HateCheck: Functional Tests for Hate Speech Detection Models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). Association for Computational Linguistics, Online, 41–58. https://doi.org/10.18653/v1/2021.acl-long.4 [42] Jatinderkumar R. Saini and Shraddha Vaidya. 2024. Recognizing Hate Speech on Twitter with Feature Combo. In Communication and Intelligent Systems, Harish Sharma, Vivek Shrivastava, Ashish Kumar Tripathi, and Lipo Wang (Eds.). Springer Nature Singapore, Singapore, 209–218. [43]Manuela Sanguinetti, Fabio Poletto, Cristina Bosco, Viviana Patti, and Marco Stranisci. 2018. An Italian Twitter Corpus of Hate Speech against Immigrants. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Nicoletta Calzolari, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Koiti Hasida, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis, and Takenobu Tokunaga (Eds.). European Language Resources Association (ELRA), Miyazaki, Japan. https: //aclanthology.org/L18-1443/ [44]Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2024. "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security (Salt Lake City, UT, USA) (CCS ’24). Association for Computing Machinery, New York, NY, USA, 1671–1685. https://doi.org/10.1145/3658644.3670388 [45]Xinyue Shen, Yixin Wu, Yiting Qu, Michael Backes, Savvas Zannettou, and Yang Zhang. 2025. HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns. In USENIX Security Symposium (USENIX Security). USENIX. [46] AI Singapore. 2024. SEA-LION (Southeast Asian Languages In One Network): A Family of Large Language Models for Southeast Asia. https: //github.com/aisingapore/sealion. [47]AI Singapore. 2024. SEA-LION (Southeast Asian Languages In One Network): A Family of Large Language Models for Southeast Asia. https: //github.com/aisingapore/sealion. [48]Xingwei Tan, Chen Lyu, Hafiz Muhammad Umer, Sahrish Khan, Mahathi Parvatham, Lois Arthurs, Simon Cullen, Shelley Wilson, Arshad Jhumka, and Gabriele Pergola. 2025. SafeSpeech: A Comprehensive and Interactive Tool for Analysing Sexist and Abusive Language in Conversations. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations), Nouha Dziri, Sean (Xiang) Ren, and Shizhe Diao (Eds.). Association for Computational Linguistics, Albuquerque, New Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia29 Mexico, 361–382. https://doi.org/10.18653/v1/2025.naacl-demo.31 [49]Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, Soroosh Mariooryad, Yifan Ding, Xinyang Geng, Fred Alcober, Roy Frostig, Mark Omernick, Lexi Walker, Cosmin Paduraru, Christina Sorokin, Andrea Tacchetti, Colin Gaffney, Samira Daruki, Olcan Sercinoglu, Zach Gleicher, Juliette Love, Paul Voigtlaender, Rohan Jain, Gabriela Surita, Kareem Mohamed, Rory Blevins, Junwhan Ahn, Tao Zhu, Kornraphop Kawintiranon, Orhan Firat, Yiming Gu, Yujing Zhang, Matthew Rahtz, Manaal Faruqui, Natalie Clay, Justin Gilmer, JD Co-Reyes, Ivo Penchev, Rui Zhu, Nobuyuki Morioka, Kevin Hui, Krishna Haridasan, Victor Campos, Mahdis Mahdieh, Mandy Guo, Samer Hassan, Kevin Kilgour, Arpi Vezer, Heng-Tze Cheng, Raoul de Liedekerke, Siddharth Goyal, Paul Barham, DJ Strouse, Seb Noury, Jonas Adler, Mukund Sundararajan, Sharad Vikram, Dmitry Lepikhin, Michela Paganini, Xavier Garcia, Fan Yang, Dasha Valter, Maja Trebacz, Kiran Vodrahalli, Chulayuth Asawaroengchai, Roman Ring, Norbert Kalb, Livio Baldini Soares, Siddhartha Brahma, David Steiner, Tianhe Yu, Fabian Mentzer, Antoine He, Lucas Gonzalez, Bibo Xu, Raphael Lopez Kaufman, Laurent El Shafey, Junhyuk Oh, Tom Hennigan, George van den Driessche, Seth Odoom, Mario Lucic, Becca Roelofs, Sid Lall, Amit Marathe, Betty Chan, Santiago Ontanon, Luheng He, Denis Teplyashin, Jonathan Lai, Phil Crone, Bogdan Damoc, Lewis Ho, Sebastian Riedel, Karel Lenc, Chih-Kuan Yeh, Aakanksha Chowdhery, Yang Xu, Mehran Kazemi, Ehsan Amid, Anastasia Petrushkina, Kevin Swersky, Ali Khodaei, Gowoon Chen, Chris Larkin, Mario Pinto, Geng Yan, Adria Puigdomenech Badia, Piyush Patil, Steven Hansen, Dave Orr, Sebastien M. R. Arnold, Jordan Grimstad, Andrew Dai, Sholto Douglas, Rishika Sinha, Vikas Yadav, Xi Chen, Elena Gribovskaya, Jacob Austin, Jeffrey Zhao, Kaushal Patel, Paul Komarek, Sophia Austin, Sebastian Borgeaud, Linda Friso, Abhimanyu Goyal, Ben Caine, Kris Cao, Da-Woon Chung, Matthew Lamm, Gabe Barth-Maron, Thais Kagohara, Kate Olszewska, Mia Chen, Kaushik Shivakumar, Rishabh Agarwal, Harshal Godhia, Ravi Rajwar, Javier Snaider, Xerxes Dotiwalla, Yuan Liu, Aditya Barua, Victor Ungureanu, Yuan Zhang, Bat-Orgil Batsaikhan, Mateo Wirth, James Qin, Ivo Danihelka, Tulsee Doshi, Martin Chadwick, Jilin Chen, Sanil Jain, Quoc Le, Arjun Kar, Madhu Gurumurthy, Cheng Li, Ruoxin Sang, Fangyu Liu, Lampros Lamprou, Rich Munoz, Nathan Lintz, Harsh Mehta, Heidi Howard, Malcolm Reynolds, Lora Aroyo, Quan Wang, Lorenzo Blanco, Albin Cassirer, Jordan Griffith, Dipanjan Das, Stephan Lee, Jakub Sygnowski, Zach Fisher, James Besley, Richard Powell, Zafarali Ahmed, Dominik Paulus, David Reitter, Zalan Borsos, Rishabh Joshi, Aedan Pope, Steven Hand, Vittorio Selo, Vihan Jain, Nikhil Sethi, Megha Goel, Takaki Makino, Rhys May, Zhen Yang, Johan Schalkwyk, Christina Butterfield, Anja Hauth, Alex Goldin, Will Hawkins, Evan Senter, Sergey Brin, Oliver Woodman, Marvin Ritter, Eric Noland, Minh Giang, Vijay Bolina, Lisa Lee, Tim Blyth, Ian Mackinnon, Machel Reid, Obaid Sarvana, David Silver, Alexander Chen, Lily Wang, Loren Maggiore, Oscar Chang, Nithya Attaluri, Gregory Thornton, Chung-Cheng Chiu, Oskar Bunyan, Nir Levine, Timothy Chung, Evgenii Eltyshev, Xiance Si, Timothy Lillicrap, Demetra Brady, Vaibhav Aggarwal, Boxi Wu, Yuanzhong Xu, Ross McIlroy, Kartikeya Badola, Paramjit Sandhu, Erica Moreira, Wojciech Stokowiec, Ross Hemsley, Dong Li, Alex Tudor, Pranav Shyam, Elahe Rahimtoroghi, Salem Haykal, Pablo Sprechmann, Xiang Zhou, Diana Mincu, Yujia Li, Ravi Addanki, Kalpesh Krishna, Xiao Wu, Alexandre Frechette, Matan Eyal, Allan Dafoe, Dave Lacey, Jay Whang, Thi Avrahami, Ye Zhang, Emanuel Taropa, Hanzhao Lin, Daniel Toyama, Eliza Rutherford, Motoki Sano, HyunJeong Choe, Alex Tomala, Chalence Safranek-Shrader, Nora Kassner, Mantas Pajarskas, Matt Harvey, Sean Sechrist, Meire Fortunato, Christina Lyu, Gamaleldin Elsayed, Chenkai Kuang, James Lottes, Eric Chu, Chao Jia, Chih-Wei Chen, Peter Humphreys, Kate Baumli, Connie Tao, Rajkumar Samuel, Cicero Nogueira dos Santos, Anders Andreassen, Nemanja Rakićević, Dominik Grewe, Aviral Kumar, Stephanie Winkler, Jonathan Caton, Andrew Brock, Sid Dalmia, Hannah Sheahan, Iain Barr, Yingjie Miao, Paul Natsev, Jacob Devlin, Feryal Behbahani, Flavien Prost, Yanhua Sun, Artiom Myaskovsky, Thanumalayan Sankaranarayana Pillai, Dan Hurt, Angeliki Lazaridou, Xi Xiong, Ce Zheng, Fabio Pardo, Xiaowei Li, Dan Horgan, Joe Stanton, Moran Ambar, Fei Xia, Alejandro Lince, Mingqiu Wang, Basil Mustafa, Albert Webson, Hyo Lee, Rohan Anil, Martin Wicke, Timothy Dozat, Abhishek Sinha, Enrique Piqueras, Elahe Dabir, Shyam Upadhyay, Anudhyan Boral, Lisa Anne Hendricks, Corey Fry, Josip Djolonga, Yi Su, Jake Walker, Jane Labanowski, Ronny Huang, Vedant Misra, Jeremy Chen, RJ Skerry-Ryan, Avi Singh, Shruti Rijhwani, Dian Yu, Alex Castro-Ros, Beer Changpinyo, Romina Datta, Sumit Bagri, Arnar Mar Hrafnkelsson, Marcello Maggioni, Daniel Zheng, Yury Sulsky, Shaobo Hou, Tom Le Paine, Antoine Yang, Jason Riesa, Dominika Rogozinska, Dror Marcus, Dalia El Badawy, Qiao Zhang, Luyu Wang, Helen Miller, Jeremy Greer, Lars Lowe Sjos, Azade Nova, Heiga Zen, Rahma Chaabouni, Mihaela Rosca, Jiepu Jiang, Charlie Chen, Ruibo Liu, Tara Sainath, Maxim Krikun, Alex Polozov, Jean-Baptiste Lespiau, Josh Newlan, Zeyncep Cankara, Soo Kwak, Yunhan Xu, Phil Chen, Andy Coenen, Clemens Meyer, Katerina Tsihlas, Ada Ma, Juraj Gottweis, Jinwei Xing, Chenjie Gu, Jin Miao, Christian Frank, Zeynep Cankara, Sanjay Ganapathy, Ishita Dasgupta, Steph Hughes-Fitt, Heng Chen, David Reid, Keran Rong, Hongmin Fan, Joost van Amersfoort, Vincent Zhuang, Aaron Cohen, Shixiang Shane Gu, Anhad Mohananey, Anastasija Ilic, Taylor Tobin, John Wieting, Anna Bortsova, Phoebe Thacker, Emma Wang, Emily Caveness, Justin Chiu, Eren Sezener, Alex Kaskasoli, Steven Baker, Katie Millican, Mohamed Elhawaty, Kostas Aisopos, Carl Lebsack, Nathan Byrd, Hanjun Dai, Wenhao Jia, Matthew Wiethoff, Elnaz Davoodi, Albert Weston, Lakshman Yagati, Arun Ahuja, Isabel Gao, Golan Pundak, Susan Zhang, Michael Azzam, Khe Chai Sim, Sergi Caelles, James Keeling, Abhanshu Sharma, Andy Swing, YaGuang Li, Chenxi Liu, Carrie Grimes Bostock, Yamini Bansal, Zachary Nado, Ankesh Anand, Josh Lipschultz, Abhijit Karmarkar, Lev Proleev, Abe Ittycheriah, Soheil Hassas Yeganeh, George Polovets, Aleksandra Faust, Jiao Sun, Alban Rrustemi, Pen Li, Rakesh Shivanna, Jeremiah Liu, Chris Welty, Federico Lebron, Anirudh Baddepudi, Sebastian Krause, Emilio Parisotto, Radu Soricut, Zheng Xu, Dawn Bloxwich, Melvin Johnson, Behnam Neyshabur, Justin Mao-Jones, Renshen Wang, Vinay Ramasesh, Zaheer Abbas, Arthur Guez, Constant Segal, Duc Dung Nguyen, James Svensson, Le Hou, Sarah York, Kieran Milan, Sophie Bridgers, Wiktor Gworek, Marco Tagliasacchi, James Lee-Thorp, Michael Chang, Alexey Guseynov, Ale Jakse Hartman, Michael Kwong, Ruizhe Zhao, Sheleem Kashem, Elizabeth Cole, Antoine Miech, Richard Tanburn, Mary Phuong, Filip Pavetic, Sebastien Cevey, Ramona Comanescu, Richard Ives, Sherry Yang, Cosmo Du, Bo Li, Zizhao Zhang, Mariko Iinuma, Clara Huiyi Hu, Aurko Roy, Shaan Bijwadia, Zhenkai Zhu, Danilo Martins, Rachel Saputro, Anita Gergely, Steven Zheng, Dawei Jia, Ioannis Antonoglou, Adam Sadovsky, Shane Gu, Yingying Bi, Alek Andreev, Sina Samangooei, Mina Khan, Tomas Kocisky, Angelos Filos, Chintu Kumar, Colton Bishop, Adams Yu, Sarah Hodkinson, Sid Mittal, Premal Shah, Alexandre Moufarek, Yong Cheng, Adam Bloniarz, Jaehoon Lee, Pedram Pejman, Paul Michel, Stephen Spencer, Vladimir Feinberg, Xuehan Xiong, Manuscript submitted to ACM 30Ng et al. Nikolay Savinov, Charlotte Smith, Siamak Shakeri, Dustin Tran, Mary Chesus, Bernd Bohnet, George Tucker, Tamara von Glehn, Carrie Muir, Yiran Mao, Hideto Kazawa, Ambrose Slone, Kedar Soparkar, Disha Shrivastava, James Cobon-Kerr, Michael Sharman, Jay Pavagadhi, Carlos Araya, Karolis Misiunas, Nimesh Ghelani, Michael Laskin, David Barker, Qiujia Li, Anton Briukhov, Neil Houlsby, Mia Glaese, Balaji Lakshminarayanan, Nathan Schucher, Yunhao Tang, Eli Collins, Hyeontaek Lim, Fangxiaoyu Feng, Adria Recasens, Guangda Lai, Alberto Magni, Nicola De Cao, Aditya Siddhant, Zoe Ashwood, Jordi Orbay, Mostafa Dehghani, Jenny Brennan, Yifan He, Kelvin Xu, Yang Gao, Carl Saroufim, James Molloy, Xinyi Wu, Seb Arnold, Solomon Chang, Julian Schrittwieser, Elena Buchatskaya, Soroush Radpour, Martin Polacek, Skye Giordano, Ankur Bapna, Simon Tokumine, Vincent Hellendoorn, Thibault Sottiaux, Sarah Cogan, Aliaksei Severyn, Mohammad Saleh, Shantanu Thakoor, Laurent Shefey, Siyuan Qiao, Meenu Gaba, Shuo yiin Chang, Craig Swanson, Biao Zhang, Benjamin Lee, Paul Kishan Rubenstein, Gan Song, Tom Kwiatkowski, Anna Koop, Ajay Kannan, David Kao, Parker Schuh, Axel Stjerngren, Golnaz Ghiasi, Gena Gibson, Luke Vilnis, Ye Yuan, Felipe Tiengo Ferreira, Aishwarya Kamath, Ted Klimenko, Ken Franko, Kefan Xiao, Indro Bhattacharya, Miteyan Patel, Rui Wang, Alex Morris, Robin Strudel, Vivek Sharma, Peter Choy, Sayed Hadi Hashemi, Jessica Landon, Mara Finkelstein, Priya Jhakra, Justin Frye, Megan Barnes, Matthew Mauger, Dennis Daun, Khuslen Baatarsukh, Matthew Tung, Wael Farhan, Henryk Michalewski, Fabio Viola, Felix de Chaumont Quitry, Charline Le Lan, Tom Hudson, Qingze Wang, Felix Fischer, Ivy Zheng, Elspeth White, Anca Dragan, Jean baptiste Alayrac, Eric Ni, Alexander Pritzel, Adam Iwanicki, Michael Isard, Anna Bulanova, Lukas Zilka, Ethan Dyer, Devendra Sachan, Srivatsan Srinivasan, Hannah Muckenhirn, Honglong Cai, Amol Mandhane, Mukarram Tariq, Jack W. Rae, Gary Wang, Kareem Ayoub, Nicholas FitzGerald, Yao Zhao, Woohyun Han, Chris Alberti, Dan Garrette, Kashyap Krishnakumar, Mai Gimenez, Anselm Levskaya, Daniel Sohn, Josip Matak, Inaki Iturrate, Michael B. Chang, Jackie Xiang, Yuan Cao, Nishant Ranka, Geoff Brown, Adrian Hutter, Vahab Mirrokni, Nanxin Chen, Kaisheng Yao, Zoltan Egyed, Francois Galilee, Tyler Liechty, Praveen Kallakuri, Evan Palmer, Sanjay Ghemawat, Jasmine Liu, David Tao, Chloe Thornton, Tim Green, Mimi Jasarevic, Sharon Lin, Victor Cotruta, Yi-Xuan Tan, Noah Fiedel, Hongkun Yu, Ed Chi, Alexander Neitz, Jens Heitkaemper, Anu Sinha, Denny Zhou, Yi Sun, Charbel Kaed, Brice Hulse, Swaroop Mishra, Maria Georgaki, Sneha Kudugunta, Clement Farabet, Izhak Shafran, Daniel Vlasic, Anton Tsitsulin, Rajagopal Ananthanarayanan, Alen Carin, Guolong Su, Pei Sun, Shashank V, Gabriel Carvajal, Josef Broder, Iulia Comsa, Alena Repina, William Wong, Warren Weilun Chen, Peter Hawkins, Egor Filonov, Lucia Loher, Christoph Hirnschall, Weiyi Wang, Jingchen Ye, Andrea Burns, Hardie Cate, Diana Gage Wright, Federico Piccinini, Lei Zhang, Chu-Cheng Lin, Ionel Gog, Yana Kulizhskaya, Ashwin Sreevatsa, Shuang Song, Luis C. Cobo, Anand Iyer, Chetan Tekur, Guillermo Garrido, Zhuyun Xiao, Rupert Kemp, Huaixiu Steven Zheng, Hui Li, Ananth Agarwal, Christel Ngani, Kati Goshvadi, Rebeca Santamaria-Fernandez, Wojciech Fica, Xinyun Chen, Chris Gorgolewski, Sean Sun, Roopal Garg, Xinyu Ye, S. M. Ali Eslami, Nan Hua, Jon Simon, Pratik Joshi, Yelin Kim, Ian Tenney, Sahitya Potluri, Lam Nguyen Thiet, Quan Yuan, Florian Luisier, Alexandra Chronopoulou, Salvatore Scellato, Praveen Srinivasan, Minmin Chen, Vinod Koverkathu, Valentin Dalibard, Yaming Xu, Brennan Saeta, Keith Anderson, Thibault Sellam, Nick Fernando, Fantine Huot, Junehyuk Jung, Mani Varadarajan, Michael Quinn, Amit Raul, Maigo Le, Ruslan Habalov, Jon Clark, Komal Jalan, Kalesha Bullard, Achintya Singhal, Thang Luong, Boyu Wang, Sujeevan Rajayogam, Julian Eisenschlos, Johnson Jia, Daniel Finchelstein, Alex Yakubovich, Daniel Balle, Michael Fink, Sameer Agarwal, Jing Li, Dj Dvijotham, Shalini Pal, Kai Kang, Jaclyn Konzelmann, Jennifer Beattie, Olivier Dousse, Diane Wu, Remi Crocker, Chen Elkind, Siddhartha Reddy Jonnalagadda, Jong Lee, Dan Holtmann-Rice, Krystal Kallarackal, Rosanne Liu, Denis Vnukov, Neera Vats, Luca Invernizzi, Mohsen Jafari, Huanjie Zhou, Lilly Taylor, Jennifer Prendki, Marcus Wu, Tom Eccles, Tianqi Liu, Kavya Kopparapu, Francoise Beaufays, Christof Angermueller, Andreea Marzoca, Shourya Sarcar, Hilal Dib, Jeff Stanway, Frank Perbet, Nejc Trdin, Rachel Sterneck, Andrey Khorlin, Dinghua Li, Xihui Wu, Sonam Goenka, David Madras, Sasha Goldshtein, Willi Gierke, Tong Zhou, Yaxin Liu, Yannie Liang, Anais White, Yunjie Li, Shreya Singh, Sanaz Bahargam, Mark Epstein, Sujoy Basu, Li Lao, Adnan Ozturel, Carl Crous, Alex Zhai, Han Lu, Zora Tung, Neeraj Gaur, Alanna Walton, Lucas Dixon, Ming Zhang, Amir Globerson, Grant Uy, Andrew Bolt, Olivia Wiles, Milad Nasr, Ilia Shumailov, Marco Selvi, Francesco Piccinno, Ricardo Aguilar, Sara McCarthy, Misha Khalman, Mrinal Shukla, Vlado Galic, John Carpenter, Kevin Villela, Haibin Zhang, Harry Richardson, James Martens, Matko Bosnjak, Shreyas Rammohan Belle, Jeff Seibert, Mahmoud Alnahlawi, Brian McWilliams, Sankalp Singh, Annie Louis, Wen Ding, Dan Popovici, Lenin Simicich, Laura Knight, Pulkit Mehta, Nishesh Gupta, Chongyang Shi, Saaber Fatehi, Jovana Mitrovic, Alex Grills, Joseph Pagadora, Dessie Petrova, Danielle Eisenbud, Zhishuai Zhang, Damion Yates, Bhavishya Mittal, Nilesh Tripuraneni, Yannis Assael, Thomas Brovelli, Prateek Jain, Mihajlo Velimirovic, Canfer Akbulut, Jiaqi Mu, Wolfgang Macherey, Ravin Kumar, Jun Xu, Haroon Qureshi, Gheorghe Comanici, Jeremy Wiesner, Zhitao Gong, Anton Ruddock, Matthias Bauer, Nick Felt, Anirudh GP, Anurag Arnab, Dustin Zelle, Jonas Rothfuss, Bill Rosgen, Ashish Shenoy, Bryan Seybold, Xinjian Li, Jayaram Mudigonda, Goker Erdogan, Jiawei Xia, Jiri Simsa, Andrea Michi, Yi Yao, Christopher Yew, Steven Kan, Isaac Caswell, Carey Radebaugh, Andre Elisseeff, Pedro Valenzuela, Kay McKinney, Kim Paterson, Albert Cui, Eri Latorre-Chimoto, Solomon Kim, William Zeng, Ken Durden, Priya Ponnapalli, Tiberiu Sosea, Christopher A. Choquette-Choo, James Manyika, Brona Robenek, Harsha Vashisht, Sebastien Pereira, Hoi Lam, Marko Velic, Denese Owusu-Afriyie, Katherine Lee, Tolga Bolukbasi, Alicia Parrish, Shawn Lu, Jane Park, Balaji Venkatraman, Alice Talbert, Lambert Rosique, Yuchung Cheng, Andrei Sozanschi, Adam Paszke, Praveen Kumar, Jessica Austin, Lu Li, Khalid Salama, Wooyeol Kim, Nandita Dukkipati, Anthony Baryshnikov, Christos Kaplanis, XiangHai Sheng, Yuri Chervonyi, Caglar Unlu, Diego de Las Casas, Harry Askham, Kathryn Tunyasuvunakool, Felix Gimeno, Siim Poder, Chester Kwak, Matt Miecnikowski, Vahab Mirrokni, Alek Dimitriev, Aaron Parisi, Dangyi Liu, Tomy Tsai, Toby Shevlane, Christina Kouridi, Drew Garmon, Adrian Goedeckemeyer, Adam R. Brown, Anitha Vijayakumar, Ali Elqursh, Sadegh Jazayeri, Jin Huang, Sara Mc Carthy, Jay Hoover, Lucy Kim, Sandeep Kumar, Wei Chen, Courtney Biles, Garrett Bingham, Evan Rosen, Lisa Wang, Qijun Tan, David Engel, Francesco Pongetti, Dario de Cesare, Dongseong Hwang, Lily Yu, Jennifer Pullman, Srini Narayanan, Kyle Levin, Siddharth Gopal, Megan Li, Asaf Aharoni, Trieu Trinh, Jessica Lo, Norman Casagrande, Roopali Vij, Loic Matthey, Bramandia Ramadhana, Austin Matthews, CJ Carey, Matthew Johnson, Kremena Goranova, Rohin Shah, Shereen Ashraf, Kingshuk Dasgupta, Rasmus Larsen, Yicheng Wang, Manish Reddy Vuyyuru, Chong Jiang, Joana Ijazi, Kazuki Osawa, Celine Smith, Ramya Sree Boppana, Taylan Bilal, Yuma Koizumi, Ying Xu, Yasemin Altun, Nir Shabat, Ben Bariach, Alex Korchemniy, Kiam Choo, Olaf Ronneberger, Chimezie Iwuanyanwu, Shubin Zhao, David Soergel, Cho-Jui Hsieh, Irene Cai, Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia31 Shariq Iqbal, Martin Sundermeyer, Zhe Chen, Elie Bursztein, Chaitanya Malaviya, Fadi Biadsy, Prakash Shroff, Inderjit Dhillon, Tejasi Latkar, Chris Dyer, Hannah Forbes, Massimo Nicosia, Vitaly Nikolaev, Somer Greene, Marin Georgiev, Pidong Wang, Nina Martin, Hanie Sedghi, John Zhang, Praseem Banzal, Doug Fritz, Vikram Rao, Xuezhi Wang, Jiageng Zhang, Viorica Patraucean, Dayou Du, Igor Mordatch, Ivan Jurin, Lewis Liu, Ayush Dubey, Abhi Mohan, Janek Nowakowski, Vlad-Doru Ion, Nan Wei, Reiko Tojo, Maria Abi Raad, Drew A. Hudson, Vaishakh Keshava, Shubham Agrawal, Kevin Ramirez, Zhichun Wu, Hoang Nguyen, Ji Liu, Madhavi Sewak, Bryce Petrini, DongHyun Choi, Ivan Philips, Ziyue Wang, Ioana Bica, Ankush Garg, Jarek Wilkiewicz, Priyanka Agrawal, Xiaowei Li, Danhao Guo, Emily Xue, Naseer Shaik, Andrew Leach, Sadh MNM Khan, Julia Wiesinger, Sammy Jerome, Abhishek Chakladar, Alek Wenjiao Wang, Tina Ornduff, Folake Abu, Alireza Ghaffarkhah, Marcus Wainwright, Mario Cortes, Frederick Liu, Joshua Maynez, Andreas Terzis, Pouya Samangouei, Riham Mansour, Tomasz Kępa, François-Xavier Aubet, Anton Algymr, Dan Banica, Agoston Weisz, Andras Orban, Alexandre Senges, Ewa Andrejczuk, Mark Geller, Niccolo Dal Santo, Valentin Anklin, Majd Al Merey, Martin Baeuml, Trevor Strohman, Junwen Bai, Slav Petrov, Yonghui Wu, Demis Hassabis, Koray Kavukcuoglu, Jeffrey Dean, and Oriol Vinyals. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv:2403.05530 [cs.CL] https://arxiv.org/abs/2403.05530 [50]Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan, Sammy Jerome, Anton Tsitsulin, Nino Vieillard, Piotr Stanczyk, Sertan Girgin, Nikola Momchev, Matt Hoffman, Shantanu Thakoor, Jean-Bastien Grill, Behnam Neyshabur, Olivier Bachem, Alanna Walton, Aliaksei Severyn, Alicia Parrish, Aliya Ahmad, Allen Hutchison, Alvin Abdagic, Amanda Carl, Amy Shen, Andy Brock, Andy Coenen, Anthony Laforge, Antonia Paterson, Ben Bastian, Bilal Piot, Bo Wu, Brandon Royal, Charlie Chen, Chintu Kumar, Chris Perry, Chris Welty, Christopher A. Choquette-Choo, Danila Sinopalnikov, David Weinberger, Dimple Vijaykumar, Dominika Rogozińska, Dustin Herbison, Elisa Bandy, Emma Wang, Eric Noland, Erica Moreira, Evan Senter, Evgenii Eltyshev, Francesco Visin, Gabriel Rasskin, Gary Wei, Glenn Cameron, Gus Martins, Hadi Hashemi, Hanna Klimczak-Plucińska, Harleen Batra, Harsh Dhand, Ivan Nardini, Jacinda Mein, Jack Zhou, James Svensson, Jeff Stanway, Jetha Chan, Jin Peng Zhou, Joana Carrasqueira, Joana Iljazi, Jocelyn Becker, Joe Fernandez, Joost van Amersfoort, Josh Gordon, Josh Lipschultz, Josh Newlan, Ju yeong Ji, Kareem Mohamed, Kartikeya Badola, Kat Black, Katie Millican, Keelin McDonell, Kelvin Nguyen, Kiranbir Sodhia, Kish Greene, Lars Lowe Sjoesund, Lauren Usui, Laurent Sifre, Lena Heuermann, Leticia Lago, Lilly McNealus, Livio Baldini Soares, Logan Kilpatrick, Lucas Dixon, Luciano Martins, Machel Reid, Manvinder Singh, Mark Iverson, Martin Görner, Mat Velloso, Mateo Wirth, Matt Davidow, Matt Miller, Matthew Rahtz, Matthew Watson, Meg Risdal, Mehran Kazemi, Michael Moynihan, Ming Zhang, Minsuk Kahng, Minwoo Park, Mofi Rahman, Mohit Khatwani, Natalie Dao, Nenshad Bardoliwalla, Nesh Devanathan, Neta Dumai, Nilay Chauhan, Oscar Wahltinez, Pankil Botarda, Parker Barnes, Paul Barham, Paul Michel, Pengchong Jin, Petko Georgiev, Phil Culliton, Pradeep Kuppala, Ramona Comanescu, Ramona Merhej, Reena Jana, Reza Ardeshir Rokni, Rishabh Agarwal, Ryan Mullins, Samaneh Saadat, Sara Mc Carthy, Sarah Cogan, Sarah Perrin, Sébastien M. R. Arnold, Sebastian Krause, Shengyang Dai, Shruti Garg, Shruti Sheth, Sue Ronstrom, Susan Chan, Timothy Jordan, Ting Yu, Tom Eccles, Tom Hennigan, Tomas Kocisky, Tulsee Doshi, Vihan Jain, Vikas Yadav, Vilobh Meshram, Vishal Dharmadhikari, Warren Barkley, Wei Wei, Wenming Ye, Woohyun Han, Woosuk Kwon, Xiang Xu, Zhe Shen, Zhitao Gong, Zichuan Wei, Victor Cotruta, Phoebe Kirk, Anand Rao, Minh Giang, Ludovic Peran, Tris Warkentin, Eli Collins, Joelle Barral, Zoubin Ghahramani, Raia Hadsell, D. Sculley, Jeanine Banks, Anca Dragan, Slav Petrov, Oriol Vinyals, Jeff Dean, Demis Hassabis, Koray Kavukcuoglu, Clement Farabet, Elena Buchatskaya, Sebastian Borgeaud, Noah Fiedel, Armand Joulin, Kathleen Kenealy, Robert Dadashi, and Alek Andreev. 2024. Gemma 2: Improving Open Language Models at a Practical Size. arXiv:2408.00118 [cs.CL] https://arxiv.org/abs/2408.00118 [51]Manuel Tonneau, Diyi Liu, Samuel Fraiberger, Ralph Schroeder, Scott A. Hale, and Paul Röttger. 2024. From Languages to Geographies: Towards Evaluating Cultural Bias in Hate Speech Datasets. In Proceedings of the 8th Workshop on Online Abuse and Harms (WOAH 2024), Yi-Ling Chung, Zeerak Talat, Debora Nozza, Flor Miriam Plaza-del Arco, Paul Röttger, Aida Mostafazadeh Davani, and Agostina Calabrese (Eds.). Association for Computational Linguistics, Mexico City, Mexico, 283–311. https://doi.org/10.18653/v1/2024.woah-1.23 [52]Zeerak Waseem and Dirk Hovy. 2016. Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter. In Proceedings of the NAACL Student Research Workshop, Jacob Andreas, Eunsol Choi, and Angeliki Lazaridou (Eds.). Association for Computational Linguistics, San Diego, California, 88–93. https://doi.org/10.18653/v1/N16-2013 [53] Zheng Weihua, Roy Ka-Wei Lee, Zhengyuan Liu, Wu Kui, AiTi Aw, and Bowei Zou. 2025. CCL-XCoT: An Efficient Cross-Lingual Knowledge Transfer Method for Mitigating Hallucination Generation. In Findings of the Association for Computational Linguistics: EMNLP 2025, Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (Eds.). Association for Computational Linguistics, Suzhou, China, 1768–1788. https://doi.org/10.18653/v1/2025.findings-emnlp.93 [54]Yunze Xiao, Yujia Hu, Kenny Tsu Wei Choo, and Roy Ka-Wei Lee. 2024. ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 6012–6025. https://doi.org/10.18653/v1/ 2024.emnlp-main.345 [55]An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfeng Xue, Na Ni, Pei Zhang, Peng Wang, Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Xuejing Liu, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, Zhifang Guo, and Zhihao Fan. 2024. Qwen2 Technical Report. arXiv:2407.10671 [cs.CL] https://arxiv.org/abs/2407.10671 Manuscript submitted to ACM 32Ng et al. CountryLegislation/Regulation Consulted IndonesiaUndang-undang Nomor 1 Tahun 2024 tentang Perubahan Kedua atas Undang- Undang Nomor 11 Tahun 2008 tentang Informasi dan Transaksi Elektronik [Law Number 1 of 2024 concerning Second Amendment to Law Number 11 of 2008 concerning Electronic Information and Transactions] 3 MalaysiaContent Code 2022 4 the Philippines The Indigenous Peoples Rights Act of 1997 5 , Safe Spaces Act 6 , Anti-Violence Against Women and Their Children Act of 2004 7 , The Revised Penal Code 8 , Magna Carta for Disabled Persons 9 SingaporeMaintenance of Religious Harmony Act 10 , the Penal Code’s section 298A 11 Thailand Thailand Civil Law Commission of Computer Related Offences Act (No. 2), BE 2560 (2017), Royal Decree on the Operation of Digital Platform Service Businesses That Are Subject to Prior Notification, B.E. 2565 (2022) Vietnam Bộ luật Lao động [Labour Code (2019)] 12 , Luật An ninh mạng [Law on Cybersecurity (2018)] 13 Table 11. Sources of Legislation consulted for defining protected categories. [56] Haotian Ye, Axel Wisiorek, Antonis Maronikolakis, Özge Alaçam, and Hinrich Schütze. 2025. A Federated Approach to Few-Shot Hate Speech Detection for Marginalized Communities. In Proceedings of the 5th Workshop on Multilingual Representation Learning (MRL 2025), David Ifeoluwa Adelani, Catherine Arnett, Duygu Ataman, Tyler A. Chang, Hila Gonen, Rahul Raja, Fabian Schmidt, David Stap, and Jiayi Wang (Eds.). Association for Computational Linguistics, Suzhuo, China, 631–651. https://doi.org/10.18653/v1/2025.mrl-main.41 [57]Xiang Yue, Yueqi Song, Akari Asai, Seungone Kim, JEAN NYANDWI, Simran Khanuja, Anjali Kantharuban, Lintang Sutawika, Sathyanarayanan Ramamoorthy, and Graham Neubig. 2025. Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages. In International Conference on Representation Learning, Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu (Eds.), Vol. 2025. 47758–47811. https://proceedings.iclr.c/paper_files/paper/ 2025/file/770b8cf7ef10b4a7170d09b36b6b6f-Paper-Conference.pdf [58]Wenxuan Zhang, Hou Pong Chan, Yiran Zhao, Mahani Aljunied, Jianyu Wang, Chaoqun Liu, Yue Deng, Zhiqiang Hu, Weiwen Xu, Yew Ken Chia, Xin Li, and Lidong Bing. 2024. SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages. arXiv:2407.19672 [cs.CL] https://arxiv.org/abs/2407.19672 [59] Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, et al.2025. MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation. arXiv preprint arXiv:2510.08608 (2025). A Data Statement A.1 Curation Rationale SEAHateCheck functional test dataset made specially to test for the socio-linguistic context of selected Southeast Asian countries, namely Indonesia, the Philippines, Thailand and Vietnam. Templates from MHC [40] and HateCheck [41] were translated by native speakers with the help of machine generated cases. In total, 31,172 test-cases were created from translated templates (SEAHateCheck Gold Label testcases) and 10926 (35%) were annotated by native speakers as hateful, non-hateful or nonsensical, and also whether more context was needed and if the testcase sounded natural. Using the Gold Label dataset as input, 19,802 SEAHateCheck Silver Label testcases were generated using LLMs, of which 400 were annotated by native speakers in the same fashion as SEAHateCheck Gold Label testcases. We also used SGHateCheck as input to generate 12,561 SGHateCheck Silver Label testcases, of which 400 were annotated by native speakers in the same fashion as SEAHateCheck Gold Label testcases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia33 Functional TestsExample No. DescriptionTemplate Label Template from SEAHateCheck Language Original English Template F1Expression of strong neg- ative emotions (explicit) HatefulI really cannot tahan TARGET like you. SinglishI really can’t stand TAR- GET like you. F2Description using very negative attributes (ex- plicit) Hateful இந்த பூமியில அடிமுட்டா ளுங்க TARGETதானு சந்ேதகேமஇல்ல TamilTARGET are without a doubt the stupidest people on this planet. F3Dehumanisation (explicit) HatefulTARGET are like can- cer. SinglishTARGET are a cancer. F4Implicit DerogationHatefulIf I wake up as a TAR- GET, I confirm suicide. SinglishIf I woke up as a TAR- GET, I would kill myself. F5Direct threatHatefulAku nak buat diorang TARGET menderita. MalayI am going to make those TARGET suffer. F6Threat as normative statement HatefulTao ước gì lũ TARGET đó phải chịu khổ. Vietnamese I would love to see those TARGET suffer. F7Hate expressed using slur Hateful ไอพวกTARGETท ํ าอ ี กแล ้ ว ThaiThe TARGET are at it again. F8Hate expressed using pro- fanity HatefulTipikal TARGET jadi orang tolol. Indonesian Typical of a TARGET to be such a total shit- head. F9Non-hateful use of profan- ity Non-HatefulTangina, ang hirap nun. TagalogFuck, that was tough. F10 Hate expressed through reference in subsequent clauses Hateful 像你这样的TARGET要知 道你的生命低贱至极。 Mandarin TARGET like you should know that your lives are worthless. F11 Hate expressed through reference in subsequent sentences HatefulGak heran sih ide-nya bullshit, soalnya dari TARGET. Indonesian I am not surprised the idea was bullshit. It came from a TARGET after all. F12 Hate expressed using negated positive state- ment HatefulTARGET nggak pantas hidup kayak manusia lain. Indonesian TARGET are not wor- thy of living like other hu- mans. F13 Non-hate expressed using negated hateful statement Non-Hateful อย ่ าย ื นเฉยด ิ พวกTARGETก ํ าล ั งโดนฆ ่ า อย ู ่ ThaiLet’s not stand by while TARGET are getting murdered. F14 Hate phrased as a ques- tion Hateful Chẳng lẽ mày không thấy, nếu lũ TARGET biến mất thì thế giới sẽ tốt đẹp hơn? Vietnamese Isn’t it apparent that we would be better off if those TARGET were eradi- cated? F15 Hate phrased as an opin- ion Hateful อยาก ให ้ ม ี คน ไป ห ั ก คอ ไอ ้ พวกTARGETจร ิ งๆ ThaiI think someone should snap the necks of those TARGET. Fig. 5. Functional Test Description and Example for F1 to F15 A.2 Language Variety SEAHateCheck Gold Label and SEAHateCheck Silver Label covers Indonesian, Tagalog, Thai and Vietnamese. SGHate- Check Silver Label covers Malay, Singlish, Tamil and Mandarin. A.3 Translator and Annotators Proficiency and Demographics All translators and annotators have the target language proficiency (Studied as a subject in school for at least 10 years and/or use it in a family setting) and use them in social situations (Read and/or write it in social media and/or use it Manuscript submitted to ACM 34Ng et al. Functional Tests Malay Singlish Tamil Mandarin Indonesian Tagalog Thai Vietnamese No. Template Label #TP #TC #TP #TC #TP #TC #TP #TC #TP #TC #TP #TC #TP #TC #TP #TC F1 Hateful 19 336 18 448 10 140 20 280 20 480 20 520 19* 456 20 600 F2 Hateful 15 210 16 308 15 210 20 280 20 480 20 520 20 480 20 600 F3 Hateful 18 246 19 296 12 146 20 280 20 412 20 424 20 416 20 532 F4 Hateful 19 281 22 462 10 140 20 280 20 480 20 520 20 480 20 600 F5 Hateful 17 223 17 299 10 140 20 280 20 446 20 472 20 448 20 566 F6 Hateful 20 336 18 360 12 168 20 280 20 480 20 520 20 480 20 600 F7 Hateful 7 50 7 33 6 18 10 60 10 40 10 70 10 240 10 70 F8 Hateful 20 328 19 350 10 118 20 280 20 463 20 496 20 464 20 583 F9 Non-Hateful 97 126 83 125 46 46 100 100 100 100 100 100 100 100 100 100 F10 Hateful 19 322 16 266 9 126 20 280 20 480 20 520 20 480 20 600 F11 Hateful 18 308 19 336 14 196 20 280 20 480 20 520 20 480 20 600 F12 Hateful 18 254 19 287 14 152 20 280 20 395 20 400 20 400 20 532 F13 Non-Hateful 20 370 15 290 12 168 20 280 20 463 20 496 20 464 20 583 F14 Hateful 19 326 14 220 12 157 20 280 20 446 20 472 20 425 20 566 F15 Hateful 18 256 17 361 13 160 20 280 20 429 20 448 20 432 20 549 F16 Non-Hateful 14 222 18 351 13 171 20 280 20 446 20 472 20 448 20 566 F17 Non-Hateful 21 406 18 448 20 269 30 420 30 703 30 756 30 704 30 883 F18 Non-Hateful 20 279 16 343 10 118 20 256 20 386 20 415 20 448 20 497 F19 Non-Hateful 20 270 16 245 9 82 20 256 20 386 20 415 20 448 20 497 F20 Non-Hateful 57 70 49 79 37 37 65 65 65 65 65 65 65 65 65 65 F21 Non-Hateful 61 74 52 84 36 36 65 65 65 65 65 65 65 65 65 65 F22 Non-Hateful 59 74 57 83 42 42 65 65 65 65 65 65 65 65 65 65 F23 Hateful - - 6 126 - - - - - - - - - - - - F24 Hateful - - 12 210 - - - - - - - - - - - - F25 Hateful - - 9 157 - - - - - - - - - - - - F26 Hateful - - 15 221 - - - - - - - - - - - - F27 Hateful - - 14 235 - - - - - - - - - - - - F32 Hateful - - - - - - 20 280 - - - - - - - - F33 Hateful - - - - - - 20 211 - - - - - - - - F34 Hateful - - - - - - 20 213 - - - - - - - - Total Number of TP and TC 596 5367 601 7023 372 2840 715 5911 655 8190 655 8751 654 8488 655 10319 Table 12. Detailed breakdown of SEAHateCheck templates (#TP) and test cases (#TC) by language and functional test. ’-’ indicates functional tests where no English templates were provided for translations. * Templates dropped as no suitable translation could be found. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia35 Functional TestsExample No. DescriptionTemplate Label Template from SEAHateCheck Language Original English Template F16 Neutral statements using protected group identifiers Non-Hateful น ี ่ เป ็ น คร ั ้ ง แรก ท ี ่ TARGETถ ู ก เสนอช ื ่ อข ึ ้ น มา ThaiThis is the first time a TARGET is nominated. F17 Positive statements using protected group identifiers Non-Hateful ก ุ ร ั กTARGETมาก ThaiI absolutely adore TAR- GET. F18 Denouncements of hate that quote it Non-HatefulUcapan kayak ”TAR- GET itu sampah” menyakitkan banget. Indonesian Statements like ”TAR- GET are scum” are deeply hurtful. F19 Denouncements of hate that make direct reference to it Non-HatefulPag sinabi mong ayaw mo talaga sa mga TAR- GET, nagpapakita lang ’yan kung gaano ka kali- gaw. TagalogSaying that you hate TARGET shows how misguided you are. F20 Abuse targeted at objects Non-HatefulAno ba ’yang mga putang inang wang-wang? TagalogWhat is it with these fuck- ing sirens? F21 Abuse targeted at individ- uals (not as member of a protected group) Non-HatefulGrabe, ang gago talaga nito TagalogHe is such a massive scum- bag. F22 Abuse targeted at non- protected groups (e.g. professions) Non-HatefulMga linta sa lipunan ang mga artist. TagalogArtists are parasites to our society. F23 Swaps of adjacent charac- ters Hateful 那种TARGET必须被处 诀。 Mandarin Those kinds of TAR- GET ought to be exe- cuetd. F24 Missing charactersHatefulTARGET are a cancr. SinglishTARGET are a cancr. F25 Missing word boundaries HatefulTARGET like you are justpathetic. SinglishTARGET like you are justpathetic. F26 Added spaces between chars Hateful 祈祷有人能纟冬结那些 TARGET。 Mandarin In a just world, those TARGET would all be d e a d. F27 Leet speak spellingsHateful 我讨yanTARGET。 Mandarin I h4te TARGET. F32 Homophone char. replace- ment HatefulWe’l tortuer these TAR- GET without mercy. SinglishWe’l tortuer these TAR- GET without mercy. F33 Character decomposition Hateful 这些TARGET就该闭上他 们的犭句嘴。 Mandarin Most TARGET are se- riously fucked in the head. F34 Pinyin spellingHateful 你和所有其他TARGET都 是社会的灾难。 Mandarin You and all other TAR- GET are a plague to so- ciety. Fig. 6. Functional Tests Description and Example for F16 to F34 with family and/or friends). Table 13 shows the dialect of the target language that the translators and student helpers speak. Before participating, all annotators were briefed about the definition of HS and protected groups in the study. We screened them on a hateful/non-hate classification task on a sample dataset, for the respective languages. In addition to the target language, all translators and annotators also met the minimum English standard to enter an English-medium university. The average age for the translators and annotators are 20.3 and 20.8 respectively. Females form a quarter of the translators and annotators. They were in their 20s and were studying for their Bachelors or recently graduated. 5 of the 8 translators and 9 of the 18 annotators WERE females. Manuscript submitted to ACM 36Ng et al. DatasetLanguageLocalities SEAHateCheck (Gold and Silver Annotators and Translators) IndonesianJakarta, Surabaya Tagalog Manila, North Luzon, Davao City ThaiBangkok Vietnamese Hanoi, Ho Chi Minh City SGHateCheck (Silver Annotators only) MalaySingapore MandarinSingapore, Malaysia Singlish Singapore TamilSingapore, Tamil Nadu Table 13. Dialect of target languages spoken by translators and annotators A.4 Data Creation Period The Indonesian templates were translated between November 2023 and February 2024. Tagalog, Thai and Vietnamese templates were translated between August 2024 and October 2024. SEAHateCheck and SGHateCheck Silver Label testcases were generated and annotated between October 2024 and January 2024. Translations were done between November 2023 and February 2024. Annotations were created between January 2024 and March 2024. A.5 Inter-annotator agreement Krippendorff ’s alpha was used to determine the inter-annotator agreement for the annotation tasks. Table 15 shows the annotation tasks in this study. A commonly accepted commonly accepted threshold for the alpha value is greater than 0.667 [24]. This was observed for all Gold Label annotations for sentiment in SEAHateCheck. For other cases, a lower alpha value was observed. HS related annotation tasks reported lower alpha values because the HS annotation task is not considered straightforward [9, 11, 35, 43]. Comparing across different annotation fields, annotators score higher agreements for task that has we observe a relatively high degree of agreement for the sentiment portion, where a majority of datasets had an alpha of above 0.667. This high score reflects the rigorous training that we gave our annotators. We speculate that the relatively lower scores for the control fields (Unnatural and Context) reflected the linguistic diversity of our annotators even within the same language. The lower scores for the additional Silver Label quality control fields (Target, Functions) reflected the complex challenge of matching the test case with the provided targets, functional tests and examples. Discussions were also held with annotators to explain the disagreements in SEAHateCheck Silver Label dataset, and can be found in Appendix B. When comparing across different datasets and languages, we can compare the relative reliability between each of them. While the SEAHateCheck Silver Label dataset have lower IAA compared to the Gold Label dataset, we observe that the control dataset, which consists of testcases where annotators had unanimous annotations, had near perfect IAA for almost all fields as expected. Hence we can attribute the lower scores of the Silver Label dataset to the increased difficulties of the task for SEAHateCheck. A different set of comparison is necessary when comparing the SGHateCheck Silver Label dataset as (1) no control fields (Unnatural, Context) were used in the original Gold Label dataset and (2) a different set of annotators annotated Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia37 Functional Tests Malay Singlish Tamil Mandarin Indonesian Tagalog Thai Vietnamese No. Template Label #TP #TC #TP #TC #TP #TC #TP #TC #TP #TC #TP #TC #TP #TC #TP #TC F1 Hateful 10 126 10 140 10 140 10 140 8 184 10 260 9* 216 10 300 F2 Hateful 8 112 8 112 15 210 8 112 9 234 8 208 8 192 8 240 F3 Hateful 10 132 11 145 12 146 10 140 12 304 10 236 10 224 10 283 F4 Hateful 10 140 12 159 10 140 10 139 11 286 10 260 10 240 10 300 F5 Hateful 10 119 10 131 10 140 10 140 10 243 10 236 10 224 10 283 F6 Hateful 10 140 10 140 12 168 10 140 9 261 10 260 10 240 10 300 F7 Hateful 4 20 4 12 6 18 4 20 4 24 4 28 4 96 4 28 F8 Hateful 10 140 10 140 10 118 10 140 10 261 10 260 10 240 10 300 F9 Non-Hateful 10 10 10 10 46 46 10 10 10 10 10 10 10 10 10 10 F10 Hateful 10 140 10 140 9 126 10 140 10 260 10 260 10 240 10 300 F11 Hateful 10 140 10 140 14 196 10 140 9 261 10 260 10 240 10 300 F12 Hateful 10 116 10 113 14 152 10 140 9 209 10 188 10 192 10 249 F13 Non-Hateful 10 132 10 131 12 168 10 140 10 243 10 236 10 224 10 283 F14 Hateful 10 124 10 122 12 157 10 140 9 253 10 212 10 208 10 266 F15 Hateful 10 132 9 117 13 160 10 140 9 243 10 236 10 224 10 283 F16 Non-Hateful 10 132 10 131 13 171 10 140 10 243 10 236 10 224 10 283 F17 Non-Hateful 10 140 10 140 20 269 10 140 10 260 10 260 10 240 10 300 F18 Non-Hateful 10 122 10 118 10 118 10 122 10 240 10 222 10 240 10 254 F19 Non-Hateful 10 106 10 100 9 82 10 122 9 186 10 174 10 208 10 220 F20 Non-Hateful 10 10 10 10 37 37 11 11 9 10 10 10 10 10 10 10 F21 Non-Hateful 10 10 10 10 36 36 10 10 9 10 10 10 10 10 10 10 F22 Non-Hateful 10 10 10 10 42 42 10 10 10 10 10 10 10 10 10 10 F23 Hateful - - 8 103 - - 5 70 - - - - - - - - F24 Hateful - - 10 131 - - - - - - - - - - - - F25 Hateful - - 10 118 - - - - - - - - - - - - F26 Hateful - - 8 92 - - 4 35 - - - - - - - - F27 Hateful - - 10 100 - - 3 32 - - - - - - - - F32 Hateful - - 3 61 - - 9 126 - - - - - - - - F33 Hateful - - 3 56 - - 10 110 - - - - - - - - F34 Hateful - - 2 42 - - 9 99 - - - - - - - - Total Number of Templates 596 5367 601 7023 372 2840 715 5911 655 8190 655 8751 654 8488 655 10319 Table 14. Breakdown of all templates and test cases generated using the templating method into their respective functional tests. #TP refers to the number of templates, #TC refers to the number of test cases. ’-’ indicates functional tests where no English templates were provided for translations. * Templates dropped as no suitable translationcould be found. Manuscript submitted to ACM 38Ng et al. Krippendorf’s Alpha DatasetLanguage SentimentUnnaturalContextTargetFunction Test Indonesian0.9020.126-0.007NANA Tagalog0.7690.0490.082NANA Thai 0.8010.1110.047NANA SEAHateCheck Gold Label Vietnamese0.8710.0650.050NANA Indonesian0.7420.1390.1520.5750.574 Tagalog0.7380.3490.0000.4100.480 Thai0.7380.3490.0000.4100.480 SEAHateCheck Silver Label Vietnamese 0.714-0.003-0.0070.2530.469 Indonesian1.0000.0001.0001.0000.824 Tagalog 1.0001.0001.0001.0000.811 Thai1.0001.0001.0001.0000.811 SEAHateCheck Silver Label (Control) Vietnamese 0.9421.0001.0000.3180.911 Malay0.6130.1840.2000.5620.401 Singlish 0.852-0.0350.3220.1890.674 Mandarin0.5530.018-0.011-0.0110.361 SGHateCheck Silver Label Tamil0.6960.0460.1980.4630.249 Malay1.000-0.023-0.0110.0000.680 Singlish0.8421.0001.000-0.0110.736 Mandarin0.490-0.079-0.012-0.0380.491 SGHateCheck Silver Label (Control) Tamil0.9120.0000.0000.0000.287 Table 15. Inter-annotator agreement for each dataset and annotation field (See section X). Cells are coloured ingreenhave a Krippendorf’s Alpha score of 1, while those inyellow have a score of 0. the Silver and Gold dataset. Hence, we should not expect almost perfect IAA as observed in SEAHateCheck Silver Label Control dataset. That said, there are some cases where the Silver Label cases have a higher IAA, particularly for Singlish. Extra caution should be placed when intepreting such results. B SEAHateCheck Gold Label Annotation Discussion To give the quantitative side of the annotation a qualitative perspective, discussions on the annotations were held after SEAHateCheck Gold Label annotations were completed. Specifically, annotated test cases were grouped according to their template and placeholders, and groups with high rate of annotation disagreement or being flagged for not being natural or requires context. Through this discussion, aim to identify whether the anomaly stems from wrongful annotation (false negatives), grammatical error or socio-cultural anomalies. The substantial findings of the annotator’s opinions where their annotations flag out quality control issues are summarised in the subsections below, and specific examples are released in with the dataset. B.1 Indonesian Testcases that sounded unnatural are usually unnecessarily long and were often a word by word translation from the English examples. Some sentences also do not really match the intended informal tone of the task. As the annotators were from two different cities in Indonesia, there were a a couple of cases where word usagwas different (e.g..dan’ is seen as too formal in Jakarta but not so in Surabaya). Outside of language issues, certain insults do not translate. There Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia39 are also certain sociocultural sensitivities that result in different sentiments for different target groups (e.g. templates relating to government). B.2 Tagalog The tone used in the translated templates sometimes shifted the sentiment of the templates. Additionally, certain phrases in Tagalog are inherently ambiguous in tone. These cases would require more context in order to be deciphered. Annotators commented that templates with double negatives tend to sound less natural. B.3 Thai The tone used in the translated templates sometimes does not match the intended sentiment of the samples. For instance, the tone of a non-hateful sentence could be really aggressive. Some of the slur placeholders were also not found to be rude enough to carry the hateful meaning to the sentences. Some translations were found to be incorrect and replacements were recommended. B.4 Vietnamese Some noun classifiers that come before placeholders do not flow. Certain placeholders were found to be unsuitable for some templates for the purpose of generating hate speech (e.g. Tao thà chết chứ không bao giờ hẹn hò với IDENTITY_S đâu. (translation: ’I’d rather die than date IDENTITY_S.’ is unsuitable for placeholders related to children). C Fine-tuning and Evaluation Prompts This section reproduces the exact prompt templates used for evaluation and fine-tuning, including label space and per-language protected-group definitions in Table 17 - 21. It shows small but important instruction differences across languages, for example Tagalog includes HIV status among protected attributes, and cites legal sources used to align operational definitions. Making these prompts public strengthens reproducibility and facilitates future baselines on SEAHateCheck. ### Instructions: Hate speech is defined as abuse targeted at a protected group or at its members for being a part of that group. Protected groups are Race, Religion, Ethnicity, Gender and Sexual Identity, Disability and Age. Using the above definition, return either "Hateful" for text classified as hate speech or "Non-Hateful" for text not classified as hate speech. Do not explain or add additional examples. Only enter Hateful or Non-Hateful and end your response. ### Text: text ### Label: label Table 17. Indonesian Manuscript submitted to ACM 40Ng et al. Open-Source Models Model NameBase Model/SOTAParamsRelease DateNotes Llama 3.2 3B Instruct (Llama3b) [14] SOTA3B25-Sep-24 Llama 3.1 8B Instruct (Llama8b) [14] SOTA8B23-Jul-24 Ministral-8B-Instruct-2410 (Ministral) [1] SOTA8B16-Oct-24 SeaLLMs-v3-7B-Chat (Seallm) [58] Qwen2 7B [55]7B30-Jul-24 Llama3.1 8B CPT SEA-LIONv3 Instruct (Sealion) [46] Llama 3.1 8B Instruct [14] 8B31-Oct-24 Gemma-9B-it (Gemma) [50]SOTA9B27-Jun-2024 Gemma2 9B CPT SEA-LIONv3 Instruct (Seagem) [47] Gemma2 9B Instruct [50] 9B19-Dec-24 Pangea-7B (Pangea) [57]Qwen2 7B [55]7B3-Oct-24 Trained on all evaluated lan- guages except Tagalog Qwen2.5-7B-Instruct (Qwen) [38] SOTA7B19-Sep-2024 Closed-Source Models Model NameBase Model/SOTAParamsRelease DateNotes o3-mini (o3) [33]SOTA-31-Jan-2025 Gemini-2.0-flash (Gemini) [10]SOTA-11-Dec-2024 deepseek-chat (Deepseek) [34]SOTA-31-Jan-2025 Table 16. Characteristics of all models tested, divided into open-source and closed-source categories. Abbreviated model names are provided in parentheses next to the full names. 3 https://peraturan.bpk.go.id/Details/274494/u-no-1-tahun-2024 4 https://w.mcmc.gov.my/skmmgovmy/media/General/registers/Content-Code-2022.pdf 5 https://elibrary.judiciary.gov.ph/thebookshelf/showdocs/2/2562 6 https://elibrary.judiciary.gov.ph/thebookshelf/showdocs/2/90094 7 https://elibrary.judiciary.gov.ph/thebookshelf/showdocs/2/22128 8 https://elibrary.judiciary.gov.ph/thebookshelf/showdocs/28/20426 9 https://elibrary.judiciary.gov.ph/thebookshelf/showdocs/2/3140 10 https://sso.agc.gov.sg/Act/MRHA1990 11 https://sso.agc.gov.sg/Act/PC1871 12 https://congbao.chinhphu.vn/noi-dung-van-ban-so-45-2019-qh14-30232 13 https://congbao.chinhphu.vn/noi-dung-van-ban-so-24-2018-qh14-26894 Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia41 ### Instructions: Hate speech is defined as abuse targeted at a protected group or at its members for being a part of that group. Protected groups are Race, Religion, Ethnicity, Gender and Sexual Identity, HIV Status, Disability and Age. Using the above definition, return either "Hateful" for text classified as hate speech or "Non-Hateful" for text not classified as hate speech. Do not explain or add additional examples. Only enter Hateful or Non-Hateful and end your response. ### Text: text ### Label: label Table 18. Tagalog ### Instructions: Hate speech is defined as abuse targeted at a protected group or at its members for being a part of that group. Protected groups are Religion, Ethnicity, Gender and Sexual Identity, Disability and Age. Using the above definition, return either "Hateful" for text classified as hate speech or "Non-Hateful" for text not classified as hate speech. Do not explain or add additional examples. Only enter Hateful or Non-Hateful and end your response. ### Text: text ### Label: label Table 19. Thai ### Instructions: Hate speech is defined as abuse targeted at a protected group or at its members for being a part of that group. Protected groups are Religion, Ethnicity, Gender, HIV Status and Sexual Identity, Disability and Age. Using the above definition, return either "Hateful" for text classified as hate speech or "Non-Hateful" for text not classified as hate speech. Do not explain or add additional examples. Only enter Hateful or Non-Hateful and end your response. ### Text: text ### Label: label Table 20. Vietnamese Manuscript submitted to ACM 42Ng et al. ### Instructions: Hate speech in Singapore is defined as abuse targeted at a protected group or at its members for being a part of that group. Protected groups are Race, Religion, Country of Origin, Gender and Sexual Identity, Disability, and Age. Using the above definition, return either "Hateful" for text classified as hate speech or "Non-Hateful" for text not classified as hate speech. Do not explain or add additional examples. Only enter Hateful or Non-Hateful and end your response. ### Text: text ### Label: label Table 21. All Singaporean Languages Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia43 Func. ClassFunctionalityG. Label # of Annotated Cases IDTGTHVNSSMSZHTA Derogation F1: Expression of strong nega- tive emotions (explicit) H----140126140140 F2: Description using very neg- ative attributes (explicit) H----84112112210 F3: Dehumanisation (explicit) (explicit) H----131132126146 F4: Implicit derogationH----303140139140 Threat. language F5: Direct threat (explicit)H----131119140140 F6: Threat as normative state- ment H----140140140168 SlursF7: Hate expressed using slurH----12201618 Profanity usage F8: Hate expressed using pro- fanity H----140140140118 F9: Non-hateful use of profan- ity NH----10101046 Pronoun refer- ence F10: Hate expressed through reference in subsequent clauses H----140140140126 F11: Hate expressed through reference in subsequent sen- tences NH----140140140196 Negation F12:Hate expressed using negated positive statement H----113116140152 F13: Non-hate expressed using negated hateful statement NH----131132140168 Phrasing F14: Hate phrased as a questionH----122124140157 F15: Hate phrased as an opin- ion H ----117132140160 Non-hateful group identifier F16: Neutral statements using protected group identifiers NH ----131132140171 F17: Positive statements using protected group identifiers NH ----140140140269 Counter speech F18: Denouncements of hate that quote it NH----118122120118 F19: Denouncements of hate that make direct reference to it NH----10010636282 Abuse against non-protected targets F20: Abuse targeted at objectsNH----10101037 F21: Abuse targeted at individ- uals (not as member of a pro- tected group) NH----10101036 F22: Abuse targeted at non- protected groups (e.g. profes- sions) NH----10101042 TotalNH----618656656865 H----2298155220831724 Total----2974225328482851 Table 22. Number of test-cases annotated in SGHateCheck across functionalities. Also shown in this table is the functional class which the functionalities belong to, its functionality number and gold labels. Manuscript submitted to ACM 44Ng et al. D Analysis over Functionalities for Gold Label Testcases D.1 Non-finetuned models D.2 Fine-tuned models Fig. 7. Accuracy across Functional Tests for Thai (left) and Vietnamese (right) Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia45 Fig. 8. Accuracy across Functional Tests for Malay (left) and Mandarin (right) Manuscript submitted to ACM 46Ng et al. Fig. 9. Accuracy across Functional Tests for Singlish (left) and Tamil (right) Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia47 E Analysis over Functionalities for Silver Label Testcases E.1 Non-finetuned models E.2 Fine-tuned models Fig. 10. Accuracy across Silver Functional Tests for Thai (left) and Vietnamese (right) Manuscript submitted to ACM 48Ng et al. Fig. 11. Accuracy across Silver Functional Tests for Malay (left) and Mandarin (right) Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia49 Fig. 12. Accuracy across Silver Functional Tests for Singlish (left) and Tamil (right) Manuscript submitted to ACM 50Ng et al. F Analysis over Protected Categories for Gold Label Testcases F.1 Non-finetuned models F.2 Fine-tuned models Fig. 13. F1 Score across Protected Categories for Thai (left) and Vietnamese (right) Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia51 Fig. 14. F1 Score across Protected Categories for Malay (left) and Mandarin (right) Manuscript submitted to ACM 52Ng et al. Fig. 15. F1 Score across Protected Categories for Singlish (left) and Tamil (right) Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia53 G Analysis over Protected Categories for Silver Label Testcases Fig. 16. F1 Score across Protected Categories for Silver Thai (left) and Vietnamese (right) Manuscript submitted to ACM 54Ng et al. Fig. 17. F1 Score across Protected Categories for Silver Malay (left) and Mandarin (right) Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia55 Fig. 18. F1 Score across Protected Categories for Silver Singlish (left) and Tamil (right) G.1 Non-finetuned models G.2 Fine-tuned models Manuscript submitted to ACM 56Ng et al. f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.86 0.94 0.78 0.81 0.99 0.82 0.88 0.99 1.00 0.97 0.99 0.57 2 derog_neg_attrib_h H 0.93 0.78 0.56 0.69 0.98 0.87 0.82 1.00 1.00 0.99 1.00 0.79 3 derog_dehum_h H 0.89 0.80 0.78 0.84 1.00 0.93 0.93 0.99 1.00 1.00 1.00 0.89 4 derog_impl_h H 0.85 0.78 0.36 0.55 0.92 0.64 0.68 0.82 0.89 0.87 0.92 0.55 5 threat_dir_h H 0.97 0.99 0.90 0.98 1.00 1.00 0.95 1.00 1.00 1.00 1.00 0.96 6 threat_norm_h H 0.95 0.93 0.83 0.97 1.00 0.97 0.97 1.00 1.00 0.98 1.00 0.93 7 slur_h H 0.79 1.00 0.79 0.79 1.00 0.71 0.64 1.00 1.00 1.00 1.00 1.00 8 profanity_h H 0.76 0.91 0.68 0.72 1.00 0.86 0.83 0.98 1.00 0.99 1.00 0.78 9 profanity_nh NH 0.80 0.70 0.90 0.80 0.60 0.90 1.00 1.00 1.00 1.00 1.00 1.00 10 ref_subs_clause_h H 0.96 0.92 0.75 0.80 0.97 0.90 0.81 0.92 0.96 0.87 1.00 0.63 11 ref_subs_sent_h H 0.98 0.96 0.77 0.87 1.00 0.99 0.79 1.00 1.00 0.99 1.00 0.77 12 negate_pos_h H 0.81 0.94 0.78 0.83 1.00 0.79 0.77 0.96 0.98 0.82 1.00 0.61 13 negate_neg_nh NH 0.25 0.40 0.59 0.77 0.48 0.48 0.78 0.58 0.55 0.93 1.00 1.00 14 phrase_question_h H 0.96 1.00 0.83 0.78 1.00 0.97 0.61 0.98 0.98 0.97 1.00 0.86 15 phrase_opinion_h H 0.84 0.86 0.59 0.76 0.91 0.89 0.82 0.95 0.95 0.98 1.00 0.79 16 ident_neutral_nh NH 0.95 0.55 0.86 0.96 0.62 0.78 0.88 0.94 0.96 1.00 0.97 1.00 17 ident_pos_nh NH 0.41 0.61 0.85 0.87 0.55 0.62 0.71 0.78 0.74 0.99 0.93 0.95 18 counter_quote_nh NH 0.17 0.26 0.44 0.36 0.06 0.16 0.54 0.26 0.10 0.12 0.30 0.29 19 counter_ref_nh NH 0.05 0.10 0.19 0.17 0.00 0.02 0.37 0.21 0.09 0.29 0.21 0.61 20 target_obj_nh NH 0.80 0.80 1.00 1.00 0.90 1.00 1.00 1.00 1.00 1.00 1.00 1.00 21 target_indiv_nh NH 0.67 0.67 0.83 0.83 0.17 0.67 0.83 0.33 0.33 0.67 1.00 0.83 22 target_group_nh NH 0.60 0.50 0.60 0.60 0.20 0.60 0.60 0.50 0.50 0.80 1.00 0.80 Table 23. Accuracy across non-finetuned models for different Functional Tests in Indonesian High-Quality Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia57 f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.85 0.79 0.76 0.58 0.99 0.70 0.67 0.89 0.96 0.84 0.96 0.57 2 derog_neg_attrib_h H 0.98 0.67 0.70 0.47 1.00 0.75 0.57 0.91 0.96 0.93 0.98 0.54 3 derog_dehum_h H 0.98 0.84 0.72 0.69 0.98 0.82 0.89 0.98 0.99 0.94 0.98 0.79 4 derog_impl_h H 0.79 0.82 0.59 0.57 0.97 0.69 0.68 0.79 0.82 0.84 0.88 0.38 5 threat_dir_h H 0.95 0.89 0.77 0.59 1.00 0.62 0.63 0.97 0.99 0.95 0.97 0.71 6 threat_norm_h H 0.97 0.73 0.66 0.64 1.00 0.80 0.80 1.00 1.00 0.96 0.98 0.83 7 slur_h H 0.50 0.08 0.33 0.25 1.00 0.25 0.25 0.83 0.83 0.67 0.75 0.42 8 profanity_h H 0.96 0.81 0.79 0.55 0.99 0.74 0.76 0.92 0.97 0.91 0.98 0.60 9 profanity_nh NH 1.00 1.00 1.00 1.00 0.75 1.00 1.00 1.00 1.00 1.00 1.00 1.00 10 ref_subs_clause_h H 0.90 0.78 0.87 0.72 1.00 0.75 0.76 0.86 0.97 0.87 0.98 0.69 11 ref_subs_sent_h H 0.91 0.80 0.90 0.62 1.00 0.70 0.56 0.91 0.95 0.92 0.98 0.65 12 negate_pos_h H 0.85 0.65 0.45 0.51 0.99 0.58 0.59 0.82 0.92 0.86 0.98 0.60 13 negate_neg_nh NH 0.32 0.36 0.33 0.73 0.12 0.35 0.57 0.63 0.59 0.85 0.87 0.99 14 phrase_question_h H 0.98 0.85 0.89 0.58 1.00 0.88 0.30 0.96 0.98 0.93 0.97 0.78 15 phrase_opinion_h H 0.86 0.83 0.65 0.66 0.96 0.70 0.72 0.99 0.99 0.93 0.97 0.70 16 ident_neutral_nh NH 0.77 0.77 0.95 0.99 0.23 0.67 0.77 0.94 0.94 0.96 0.90 0.99 17 ident_pos_nh NH 0.43 0.50 0.72 0.89 0.25 0.48 0.73 0.63 0.69 0.91 0.91 0.97 18 counter_quote_nh NH 0.17 0.39 0.15 0.41 0.01 0.10 0.68 0.39 0.14 0.19 0.14 0.55 19 counter_ref_nh NH 0.01 0.40 0.26 0.55 0.10 0.13 0.70 0.39 0.15 0.40 0.29 0.86 20 target_obj_nh NH 0.67 0.67 0.83 0.83 0.33 0.83 1.00 0.83 0.83 1.00 1.00 1.00 21 target_indiv_nh NH 0.71 0.71 0.86 1.00 0.43 0.86 1.00 0.71 0.71 1.00 1.00 1.00 22 target_group_nh NH 1.00 1.00 1.00 1.00 0.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 Table 24. Accuracy across non-finetuned models for different Functional Tests in Tagalog High-Quality Test Cases. Manuscript submitted to ACM 58Ng et al. f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.93 0.95 0.94 0.85 0.94 0.69 0.88 0.94 0.96 0.90 0.91 0.65 2 derog_neg_attrib_h H 0.99 0.98 0.97 0.84 0.94 0.48 0.95 0.95 0.97 0.93 0.91 0.67 3 derog_dehum_h H 1.00 0.83 0.92 0.84 0.94 0.63 0.94 0.95 0.96 0.92 0.91 0.72 4 derog_impl_h H 0.91 0.93 0.80 0.77 0.89 0.66 0.83 0.83 0.83 0.88 0.85 0.43 5 threat_dir_h H 6 threat_norm_h H 7 slur_h H 0.70 0.52 0.70 0.56 0.73 0.29 0.49 0.55 0.62 0.78 0.82 0.15 8 profanity_h H 0.84 0.79 0.90 0.78 0.99 0.65 0.83 0.89 0.92 0.88 0.90 0.66 9 profanity_nh NH 0.60 0.40 0.80 1.00 0.80 1.00 1.00 1.00 1.00 1.00 1.00 1.00 10 ref_subs_clause_h H 0.94 0.91 0.89 0.91 1.00 0.92 0.99 0.99 0.99 0.94 0.92 0.75 11 ref_subs_sent_h H 0.94 0.85 0.95 0.81 0.99 0.71 0.91 0.95 0.97 0.93 0.91 0.68 12 negate_pos_h H 0.98 0.88 0.91 0.85 0.89 0.48 0.89 0.91 0.93 0.91 0.90 0.64 13 negate_neg_nh NH 0.40 0.59 0.54 0.78 0.57 0.66 0.74 0.69 0.69 0.79 0.99 0.99 14 phrase_question_h H 0.98 0.92 0.98 0.92 0.96 0.75 0.96 0.92 0.95 0.91 0.91 0.70 15 phrase_opinion_h H 0.97 1.00 0.92 0.83 0.99 0.74 0.95 0.83 0.87 0.92 0.87 0.68 16 ident_neutral_nh NH 0.81 0.73 0.77 0.92 0.77 0.86 0.82 0.80 0.79 0.79 0.86 0.98 17 ident_pos_nh NH 0.60 0.77 0.72 0.88 0.69 0.82 0.76 0.85 0.85 0.94 0.98 0.97 18 counter_quote_nh NH 0.08 0.32 0.10 0.14 0.03 0.09 0.21 0.30 0.11 0.17 0.32 0.52 19 counter_ref_nh NH 0.07 0.19 0.07 0.25 0.06 0.24 0.16 0.36 0.21 0.37 0.57 0.71 20 target_obj_nh NH 0.96 0.89 0.95 0.93 0.95 0.71 0.96 0.96 0.98 0.97 0.92 0.82 21 target_indiv_nh NH 1.00 1.00 0.94 0.97 0.94 0.77 0.93 1.00 1.00 0.93 0.90 0.88 22 target_group_nh NH 0.83 0.67 1.00 1.00 0.67 1.00 0.83 0.67 0.83 0.83 1.00 1.00 Table 25. Accuracy across non-finetuned models for different Functional Tests in Thai High-Quality Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia59 f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.99 0.97 0.70 0.71 1.00 0.94 0.96 0.96 0.97 0.92 0.976666667 0.78 2 derog_neg_attrib_h H 0.97 0.93 0.84 0.71 0.99 0.90 0.97 0.96 0.96 0.93 1.00 0.81 3 derog_dehum_h H 0.87 0.83 0.76 0.83 0.99 0.87 0.98 0.96 0.97 0.96 0.98 0.88 4 derog_impl_h H 0.92 0.84 0.59 0.50 1.00 0.84 0.73 0.87 0.91 0.81 0.91 0.57 5 threat_dir_h H 0.99 0.99 0.96 0.98 1.00 0.96 0.95 1.00 1.00 0.95 1 0.93 6 threat_norm_h H 7 slur_h H 0.89 0.67 0.89 0.44 1.00 0.78 0.44 0.67 0.89 0.67 0.33 0.33 8 profanity_h H 0.96 0.85 0.83 0.81 1.00 0.84 0.91 0.97 0.97 0.94 0.97 0.75 9 profanity_nh NH 0.63 0.75 1.00 1.00 0.25 1.00 1.00 0.88 0.88 1 0.875 1 10 ref_subs_clause_h H 0.94 0.86 0.95 0.73 1.00 0.96 0.96 0.98 0.99 0.91 0.99 0.77 11 ref_subs_sent_h H 0.99 0.89 0.92 0.87 1.00 0.96 0.94 0.99 0.99 0.95 0.98 0.77 12 negate_pos_h H 0.91 0.94 0.79 0.83 0.99 0.75 0.93 0.94 0.95 0.86 0.98 0.77 13 negate_neg_nh NH 0.59 0.54 0.59 0.85 0.71 0.85 0.78 0.91 0.88 0.82 1 1 14 phrase_question_h H 0.96 0.92 0.85 0.84 0.98 0.88 0.77 0.93 0.96 0.92 0.99 0.78 15 phrase_opinion_h H 0.99 0.96 0.87 0.80 0.99 0.92 0.89 0.97 0.98 0.90 0.99 0.83 16 ident_neutral_nh NH 0.82 0.79 0.92 1.00 0.50 0.64 0.77 0.91 0.89 0.90 0.93 1 17 ident_pos_nh NH 0.59 0.69 0.87 0.94 0.73 0.76 0.77 0.83 0.83 0.95 0.97 0.99 18 counter_quote_nh NH 0.23 0.28 0.22 0.28 0.00 0.02 0.30 0.17 0.07 0.11 0.15 0.28 19 counter_ref_nh NH 0.05 0.24 0.31 0.37 0.03 0.32 0.28 0.27 0.18 0.30 0.45 0.61 20 target_obj_nh NH 1.00 1.00 0.93 0.92 1.00 0.99 0.88 1.00 1.00 0.95 0.96 0.88 21 target_indiv_nh NH 0.00 0.50 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1 1 1 22 target_group_nh NH 0.00 0.00 1.00 1.00 0.00 0.00 0.00 0.00 0.00 1 1 1 Table 26. Accuracy across non-finetuned models for different Functional Tests in Vietnamese High-Quality Test Cases. Manuscript submitted to ACM 60Ng et al. f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.82 0.91 0.88 0.74 0.98 0.85 0.85 0.95 0.98 0.99 1.00 0.71 2 derog_neg_attrib_h H 0.85 0.97 0.91 0.78 1.00 0.93 0.90 0.98 0.98 1.00 1.00 0.88 3 derog_dehum_h H 0.89 0.84 0.92 0.86 0.98 0.85 0.89 0.98 0.98 1.00 0.99 0.90 4 derog_impl_h H 0.84 0.84 0.78 0.70 0.98 0.84 0.79 0.88 0.91 0.93 0.87 0.64 5 threat_dir_h H 0.91 0.97 0.97 0.93 1.00 0.92 0.96 0.98 0.98 0.99 0.97 0.92 6 threat_norm_h H 0.97 0.96 0.95 0.87 0.99 0.98 0.97 1.00 1.00 1.00 1.00 0.94 7 slur_h H 0.75 0.75 1.00 0.50 1.00 1.00 1.00 0.75 1.00 0.75 1.00 0.00 8 profanity_h H 0.86 0.86 0.87 0.71 0.99 0.87 0.85 0.94 0.96 1.00 1.00 0.83 9 profanity_nh NH 0.71 0.43 0.86 1.00 0.29 1.00 0.86 0.86 0.86 1.00 1.00 1.00 10 ref_subs_clause_h H 0.88 0.89 0.77 0.77 0.98 0.92 0.86 0.90 0.92 0.95 0.99 0.73 11 ref_subs_sent_h H 0.86 0.86 0.78 0.74 1.00 0.95 0.82 0.96 0.96 1.00 1.00 0.76 12 negate_pos_h H 0.79 0.88 0.74 0.70 0.93 0.78 0.88 0.93 0.94 0.99 0.99 0.74 13 negate_neg_nh NH 0.39 0.27 0.39 0.83 0.39 0.39 0.57 0.66 0.68 0.69 0.95 0.99 14 phrase_question_h H 0.90 0.97 0.97 0.80 0.99 0.97 0.80 0.95 0.97 0.97 0.98 0.90 15 phrase_opinion_h H 0.95 1.00 0.81 0.81 1.00 0.96 0.92 0.96 0.98 1.00 0.99 0.93 16 ident_neutral_nh NH 0.85 0.53 0.75 0.97 0.42 0.63 0.63 0.87 0.89 0.86 0.94 0.98 17 ident_pos_nh NH 0.43 0.69 0.66 0.87 0.54 0.61 0.61 0.66 0.71 0.89 0.87 0.98 18 counter_quote_nh NH 0.14 0.04 0.10 0.24 0.00 0.11 0.30 0.19 0.03 0.03 0.17 0.23 19 counter_ref_nh NH 0.11 0.04 0.11 0.29 0.00 0.09 0.24 0.28 0.15 0.10 0.30 0.46 20 target_obj_nh NH 0.89 0.78 1.00 1.00 0.78 0.89 0.89 0.89 0.89 1.00 0.89 1.00 21 target_indiv_nh NH 0.67 0.56 0.44 0.78 0.22 0.78 0.67 0.44 0.44 0.44 1.00 0.89 22 target_group_nh NH 0.60 0.60 0.80 0.80 0.40 0.80 0.60 0.80 0.80 0.40 1.00 0.80 Table 27. Accuracy across non-finetuned models for different Functional Tests in Malay High-Quality Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia61 f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.94 0.95 0.96 0.77 1.00 0.87 0.98 0.91 0.95 0.98 1.00 0.73 2 derog_neg_attrib_h H 0.97 0.98 0.98 0.80 0.99 0.87 0.99 0.94 0.97 0.99 1.00 0.75 3 derog_dehum_h H 0.97 0.94 0.98 0.83 1.00 0.88 0.98 0.96 0.97 1.00 1.00 0.86 4 derog_impl_h H 0.96 0.78 0.87 0.54 0.93 0.74 0.95 0.80 0.88 0.98 0.96 0.71 5 threat_dir_h H 0.96 0.96 0.96 0.79 1.00 0.92 0.92 0.88 0.96 0.88 0.96 0.71 6 threat_norm_h H 0.99 0.91 1.00 0.88 1.00 0.95 0.95 0.91 0.91 1.00 1.00 0.81 7 slur_h H 0.33 0.33 0.33 0.00 1.00 0.67 0.33 1.00 1.00 1.00 0.33 0.67 8 profanity_h H 0.99 0.89 1.00 0.92 1.00 0.94 1.00 0.98 0.98 1.00 1.00 0.87 9 profanity_nh NH 1.00 1.00 0.67 1.00 0.67 0.67 1.00 1.00 1.00 1.00 1.00 1.00 10 ref_subs_clause_h H 0.99 0.97 1.00 0.92 0.99 0.99 0.95 0.96 1.00 1.00 1.00 0.83 11 ref_subs_sent_h H 0.98 0.83 0.99 0.88 0.99 0.94 0.89 0.93 0.97 1.00 1.00 0.69 12 negate_pos_h H 0.92 0.93 0.94 0.83 0.96 0.85 0.95 0.93 0.97 0.98 1.00 0.66 13 negate_neg_nh NH 0.36 0.36 0.32 0.80 0.45 0.50 0.50 0.75 0.67 0.72 0.96 0.99 14 phrase_question_h H 0.99 0.98 1.00 0.94 1.00 0.98 0.93 0.96 0.99 1.00 1.00 0.91 15 phrase_opinion_h H 0.99 1.00 0.98 0.93 1.00 0.98 0.99 0.97 0.99 0.99 1.00 0.84 16 ident_neutral_nh NH 0.92 0.58 0.75 1.00 0.58 0.83 0.75 1.00 1.00 0.92 0.75 1.00 17 ident_pos_nh NH 0.58 0.55 0.53 0.89 0.63 0.79 0.63 0.97 0.95 0.95 0.92 1.00 18 counter_quote_nh NH 0.11 0.46 0.03 0.46 0.03 0.28 0.49 0.44 0.08 0.06 0.39 0.36 19 counter_ref_nh NH 0.03 0.17 0.09 0.36 0.06 0.17 0.20 0.40 0.17 0.12 0.22 0.60 20 target_obj_nh NH 0.96 0.79 0.81 0.72 1.00 0.90 0.87 0.82 0.85 0.93 0.93 0.53 21 target_indiv_nh NH 0.95 0.78 0.97 0.78 1.00 0.92 0.98 0.95 0.95 1.00 1.00 0.83 22 target_group_nh NH 0.97 0.97 0.95 0.78 1.00 0.96 0.93 0.84 0.87 0.99 0.93 0.89 Table 28. Accuracy across non-finetuned models for different Functional Tests in Mandarin High-Quality Test Cases. Manuscript submitted to ACM 62Ng et al. f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.91 1.00 0.91 0.82 0.96 0.86 0.95 0.87 0.92 0.97 0.99 0.87 2 derog_neg_attrib_h H 0.90 0.99 0.96 0.79 0.99 0.89 0.95 0.89 0.93 0.99 0.99 0.87 3 derog_dehum_h H 0.86 0.96 0.95 0.87 0.91 0.82 0.95 0.94 0.96 0.96 0.99 0.92 4 derog_impl_h H 0.93 0.93 0.96 0.85 0.92 0.91 0.91 0.84 0.86 0.97 0.98 0.81 5 threat_dir_h H 0.93 0.99 0.97 0.83 1.00 0.92 0.94 0.91 0.94 0.98 0.99 0.84 6 threat_norm_h H 0.91 0.96 0.92 0.82 0.84 0.81 0.96 0.97 0.99 0.98 0.99 0.92 7 slur_h H 1.00 0.71 0.86 0.71 1.00 0.71 0.57 1.00 1.00 1.00 0.71 1.00 8 profanity_h H 0.86 0.97 0.96 0.86 0.92 0.85 0.94 0.89 0.95 1.00 0.99 0.91 9 profanity_nh NH 1.00 0.20 0.80 1.00 0.40 0.80 0.80 1.00 1.00 1.00 1.00 1.00 10 ref_subs_clause_h H 0.98 1.00 1.00 0.93 0.90 0.85 0.99 0.91 0.99 0.99 1.00 0.94 11 ref_subs_sent_h H 0.93 0.99 0.99 0.88 0.95 0.85 0.93 0.93 0.96 0.99 0.99 0.91 12 negate_pos_h H 0.91 0.97 0.94 0.92 0.96 0.84 0.98 0.90 0.96 0.99 1.00 0.88 13 negate_neg_nh NH 0.68 0.43 0.43 0.74 0.51 0.55 0.56 0.68 0.69 0.73 0.86 0.87 14 phrase_question_h H 0.92 1.00 1.00 0.97 0.89 0.86 0.93 0.96 0.99 0.98 1.00 0.94 15 phrase_opinion_h H 0.94 1.00 0.99 0.95 1.00 0.95 1.00 0.96 0.98 1.00 1.00 0.95 16 ident_neutral_nh NH 0.75 0.44 0.78 0.89 0.36 0.67 0.78 0.88 0.88 0.74 0.65 0.97 17 ident_pos_nh NH 0.75 0.55 0.79 0.90 0.74 0.62 0.66 0.85 0.85 0.90 0.90 0.97 18 counter_quote_nh NH 0.15 0.08 0.03 0.11 0.05 0.04 0.04 0.05 0.00 0.01 0.12 0.10 19 counter_ref_nh NH 0.11 0.12 0.06 0.22 0.15 0.11 0.17 0.22 0.11 0.11 0.17 0.40 20 target_obj_nh NH 0.89 0.99 1.00 0.91 0.85 0.79 0.98 0.95 0.96 0.99 1.00 0.91 21 target_indiv_nh NH 0.87 0.95 0.91 0.73 0.92 0.79 0.85 0.88 0.88 0.84 0.85 0.75 22 target_group_nh NH 0.89 0.98 0.98 0.91 0.98 0.93 0.98 0.95 0.96 1.00 1.00 0.93 Table 29. Accuracy across non-finetuned models for different Functional Tests in Singlish High-Quality Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia63 f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.86 0.68 1.00 0.68 0.88 0.52 0.45 0.81 0.92 1.00 0.79 0.55 2 derog_neg_attrib_h H 0.83 0.83 0.93 0.79 0.99 0.67 0.66 0.86 0.93 0.98 0.87 0.51 3 derog_dehum_h H 0.93 0.91 0.99 0.92 1.00 0.70 0.70 0.95 0.98 1.00 0.98 0.69 4 derog_impl_h H 0.66 0.62 0.77 0.56 0.94 0.47 0.52 0.69 0.73 0.93 0.79 0.29 5 threat_dir_h H 0.88 0.66 0.98 0.83 0.94 0.58 0.61 0.89 0.96 0.99 0.88 0.77 6 threat_norm_h H 0.87 0.79 0.99 0.83 0.99 0.73 0.67 0.93 0.98 1.00 0.95 0.79 7 slur_h H 0.13 0.38 1.00 0.75 1.00 0.13 0.50 0.13 0.50 0.75 0.25 0.00 8 profanity_h H 0.84 0.72 0.98 0.80 1.00 0.59 0.50 0.83 0.90 1.00 0.92 0.51 9 profanity_nh NH 1.00 0.84 0.52 1.00 0.23 0.94 0.97 1.00 1.00 0.97 0.97 1.00 10 ref_subs_clause_h H 0.76 0.70 1.00 0.90 0.99 0.63 0.55 0.88 0.98 0.98 0.90 0.56 11 ref_subs_sent_h H 0.91 0.74 0.99 0.91 1.00 0.68 0.56 0.89 0.97 0.96 0.93 0.60 12 negate_pos_h H 0.82 0.80 0.99 0.85 0.99 0.60 0.73 0.88 0.95 0.98 0.97 0.59 13 negate_neg_nh NH 0.41 0.50 0.03 0.34 0.16 0.38 0.59 0.46 0.40 0.38 0.87 0.90 14 phrase_question_h H 0.74 0.81 0.98 0.78 1.00 0.60 0.32 0.83 0.89 0.91 0.86 0.49 15 phrase_opinion_h H 0.84 0.79 0.98 0.79 1.00 0.42 0.69 0.92 0.97 0.97 0.99 0.53 16 ident_neutral_nh NH 0.69 0.82 0.39 0.88 0.22 0.65 0.57 0.96 0.98 0.72 0.98 1.00 17 ident_pos_nh NH 0.38 0.64 0.29 0.88 0.18 0.62 0.59 0.76 0.78 0.80 0.93 1.00 18 counter_quote_nh NH 0.22 0.28 0.00 0.13 0.02 0.40 0.52 0.33 0.05 0.07 0.22 0.58 19 counter_ref_nh NH 0.16 0.18 0.00 0.18 0.00 0.25 0.73 0.34 0.25 0.16 0.39 0.55 20 target_obj_nh NH 0.81 0.65 0.45 0.94 0.10 0.71 0.94 1.00 1.00 1.00 0.94 1.00 21 target_indiv_nh NH 0.30 0.47 0.10 0.77 0.33 0.83 0.63 0.73 0.67 0.50 1.00 0.90 22 target_group_nh NH 0.57 0.43 0.22 0.83 0.17 0.52 0.70 0.74 0.74 0.65 0.96 1.00 Table 30. Accuracy across non-finetuned models for different Functional Tests in Tamil High-Quality Test Cases. Manuscript submitted to ACM 64Ng et al. f_n t_function t_g Ministral Llama3b Llama8b sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.42 0.41 0.78 0.89 0.58 0.54 0.60 0.86 0.80 2 derog_neg_attrib_h H 0.54 0.47 0.69 0.88 0.86 0.64 0.82 0.84 0.69 3 derog_dehum_h H 0.62 0.56 0.86 0.89 0.83 0.76 0.90 0.87 0.81 4 derog_impl_h H 0.49 0.16 0.53 0.59 0.37 0.20 0.63 0.65 0.48 5 threat_dir_h H 0.51 0.77 0.91 0.96 0.91 0.76 0.76 0.94 0.81 6 threat_norm_h H 0.45 0.69 0.85 0.92 0.81 0.88 0.90 0.99 0.97 7 slur_h H 0.86 0.79 0.93 0.86 0.86 0.79 0.86 0.86 0.86 8 profanity_h H 0.48 0.54 0.75 0.90 0.59 0.67 0.78 0.91 0.83 9 profanity_nh NH 0.90 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 10 ref_subs_clause_h H 0.32 0.49 0.60 0.75 0.77 0.64 0.75 0.55 0.33 11 ref_subs_sent_h H 0.82 0.63 0.83 1.00 0.89 0.94 0.94 0.96 0.96 12 negate_pos_h H 0.27 0.36 0.61 0.64 0.76 0.41 0.57 0.90 0.63 13 negate_neg_nh NH 0.96 0.77 0.79 0.86 0.90 0.91 0.91 0.85 0.98 14 phrase_question_h H 0.51 0.71 0.77 0.93 0.85 0.83 0.70 0.96 0.87 15 phrase_opinion_h H 0.43 0.52 0.69 0.72 0.74 0.49 0.76 0.79 0.65 16 ident_neutral_nh NH 0.99 0.99 0.98 0.99 0.98 0.99 0.97 0.99 1.00 17 ident_pos_nh NH 0.83 0.93 0.91 0.93 0.96 0.97 0.96 0.90 1.00 18 counter_quote_nh NH 0.47 0.53 0.61 0.31 0.22 0.34 0.40 0.63 0.64 19 counter_ref_nh NH 0.50 0.32 0.24 0.26 0.21 0.31 0.23 0.66 0.69 20 target_obj_nh NH 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 21 target_indiv_nh NH 0.67 0.83 0.50 0.33 0.33 0.50 0.50 0.83 0.83 22 target_group_nh NH 0.60 0.70 0.60 0.40 0.70 0.70 0.70 0.80 0.70 Table 31. Accuracy across fine-tuned models for different Functional Tests in Indonesian High-Quality Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia65 f_n t_function t_g Ministral Llama3b Llama8b sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.83 0.88 1.00 1.00 0.90 0.94 0.88 0.95 1.00 2 derog_neg_attrib_h H 0.94 0.85 0.97 0.99 1.00 0.84 0.82 0.99 1.00 3 derog_dehum_h H 0.98 0.96 0.99 1.00 1.00 1.00 1.00 1.00 1.00 4 derog_impl_h H 0.80 0.81 0.93 0.94 0.85 0.92 0.85 0.97 0.93 5 threat_dir_h H 0.72 0.90 0.89 0.84 0.95 0.93 0.73 1.00 1.00 6 threat_norm_h H 0.86 0.87 1.00 1.00 0.97 0.92 0.90 1.00 1.00 7 slur_h H 0.83 0.50 1.00 1.00 0.75 0.67 0.75 0.83 0.83 8 profanity_h H 1.00 1.00 0.97 0.99 0.98 1.00 0.96 1.00 1.00 9 profanity_nh NH 0.50 0.50 0.25 0.50 0.50 0.50 0.00 0.50 0.75 10 ref_subs_clause_h H 0.99 1.00 1.00 1.00 0.95 0.94 0.99 1.00 1.00 11 ref_subs_sent_h H 0.91 0.99 1.00 0.99 0.96 0.95 1.00 1.00 1.00 12 negate_pos_h H 0.83 0.78 0.91 0.96 0.91 0.85 0.77 0.88 0.93 13 negate_neg_nh NH 0.41 0.17 0.25 0.49 0.24 0.21 0.26 0.63 0.68 14 phrase_question_h H 0.97 0.88 1.00 1.00 1.00 0.95 0.97 1.00 1.00 15 phrase_opinion_h H 0.93 0.97 0.97 0.97 0.90 1.00 0.94 1.00 1.00 16 ident_neutral_nh NH 0.76 0.83 0.60 0.71 0.76 0.74 0.65 0.90 0.88 17 ident_pos_nh NH 0.60 0.66 0.73 0.81 0.81 0.56 0.69 0.78 0.92 18 counter_quote_nh NH 0.08 0.08 0.00 0.07 0.02 0.14 0.27 0.25 0.18 19 counter_ref_nh NH 0.00 0.06 0.00 0.00 0.05 0.11 0.12 0.01 0.13 20 target_obj_nh NH 0.50 0.50 0.33 0.50 0.67 0.50 0.50 0.50 0.50 21 target_indiv_nh NH 0.29 0.43 0.14 0.14 0.57 0.57 0.43 0.00 0.00 22 target_group_nh NH 0.00 0.00 0.00 0.00 1.00 1.00 0.00 0.00 0.00 Table 32. Accuracy across fine-tuned models for different Functional Tests in Tagalog High-Quality Test Cases. Manuscript submitted to ACM 66Ng et al. f_n t_function t_g Ministral Llama3b Llama8b sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.61 0.50 0.95 0.82 0.80 0.58 0.64 0.64 0.64 2 derog_neg_attrib_h H 0.51 0.52 0.97 0.93 0.72 0.58 0.75 0.86 0.92 3 derog_dehum_h H 0.85 0.71 0.94 0.86 0.91 0.82 0.96 0.98 0.90 4 derog_impl_h H 0.69 0.59 0.77 0.74 0.74 0.63 0.76 0.74 0.55 5 threat_dir_h H 6 threat_norm_h H 7 slur_h H 0.10 0.12 0.48 0.33 0.22 0.05 0.16 0.27 0.20 8 profanity_h H 0.62 0.78 0.97 0.91 0.77 0.69 0.80 0.82 0.87 9 profanity_nh NH 1.00 0.80 0.80 0.80 0.80 1.00 0.80 1.00 1.00 10 ref_subs_clause_h H 0.80 0.93 0.95 0.98 0.96 0.79 0.91 0.97 0.95 11 ref_subs_sent_h H 0.56 0.81 0.74 0.75 0.75 0.80 0.88 0.96 0.96 12 negate_pos_h H 0.64 0.78 0.91 0.85 0.86 0.71 0.75 0.95 0.81 13 negate_neg_nh NH 0.96 0.84 0.75 0.75 0.93 0.86 0.92 0.98 0.99 14 phrase_question_h H 0.70 0.65 0.91 0.83 0.80 0.74 0.89 0.87 0.82 15 phrase_opinion_h H 0.66 0.85 0.88 0.87 0.88 0.75 0.87 0.87 0.87 16 ident_neutral_nh NH 1.00 0.99 0.94 0.98 0.98 0.99 1.00 1.00 1.00 17 ident_pos_nh NH 0.99 0.99 0.98 0.99 0.96 1.00 1.00 0.99 0.99 18 counter_quote_nh NH 0.47 0.41 0.18 0.16 0.24 0.32 0.47 0.66 0.44 19 counter_ref_nh NH 0.75 0.49 0.36 0.55 0.32 0.44 0.42 0.73 0.71 20 target_obj_nh NH 0.88 0.80 0.97 0.93 0.88 0.82 0.96 0.95 0.90 21 target_indiv_nh NH 0.78 0.85 0.95 0.96 0.84 0.83 0.83 0.96 0.89 22 target_group_nh NH 1.00 0.83 0.83 1.00 0.83 1.00 1.00 1.00 1.00 Table 33. Accuracy across fine-tuned models for different Functional Tests in Thai High-Quality Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia67 f_n t_function t_g Ministral Llama3b Llama8b sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.95 0.98 1.00 0.99 0.99 0.99 1.00 1.00 1.00 2 derog_neg_attrib_h H 1.00 0.91 1.00 1.00 1.00 1.00 1.00 1.00 1.00 3 derog_dehum_h H 1.00 1.00 1.00 1.00 1.00 1.00 1.00 0.99 1.00 4 derog_impl_h H 0.50 0.92 0.93 0.96 0.98 0.97 0.94 0.94 0.95 5 threat_dir_h H 0.98 0.97 1.00 1.00 1.00 1.00 1.00 1.00 1.00 6 threat_norm_h H 7 slur_h H 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 8 profanity_h H 1.00 0.98 1.00 1.00 1.00 1.00 1.00 0.99 1.00 9 profanity_nh NH 0.50 0.50 0.50 0.25 0.50 0.38 0.63 0.63 0.50 10 ref_subs_clause_h H 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 11 ref_subs_sent_h H 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 12 negate_pos_h H 0.96 0.99 1.00 1.00 0.98 0.98 0.98 0.98 0.97 13 negate_neg_nh NH 0.78 0.60 0.73 0.77 0.82 0.81 0.64 0.90 0.90 14 phrase_question_h H 0.99 0.97 0.99 1.00 1.00 0.99 1.00 1.00 1.00 15 phrase_opinion_h H 1.00 1.00 0.99 0.99 1.00 1.00 0.99 1.00 1.00 16 ident_neutral_nh NH 0.91 0.83 0.78 0.88 0.88 0.96 0.91 0.95 0.95 17 ident_pos_nh NH 0.86 0.88 0.91 0.92 0.87 0.92 0.87 0.91 0.92 18 counter_quote_nh NH 0.12 0.03 0.00 0.00 0.00 0.00 0.01 0.22 0.09 19 counter_ref_nh NH 0.11 0.09 0.03 0.10 0.04 0.04 0.02 0.55 0.16 20 target_obj_nh NH 1.00 0.98 1.00 1.00 1.00 1.00 0.99 1.00 1.00 21 target_indiv_nh NH 0.50 1.00 0.50 0.50 1.00 1.00 1.00 0.50 0.50 22 target_group_nh NH 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 Table 34. Accuracy across fine-tuned models for different Functional Tests in Vietnamese High-Quality Test Cases. Manuscript submitted to ACM 68Ng et al. f_n t_function t_g Ministral Llama3b Llama8b sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.55 0.51 0.93 0.90 0.74 0.65 0.78 0.77 0.75 2 derog_neg_attrib_h H 0.75 0.79 0.97 0.96 0.91 0.83 0.97 0.84 0.83 3 derog_dehum_h H 0.68 0.47 0.92 0.93 0.81 0.69 0.93 0.90 0.77 4 derog_impl_h H 0.51 0.44 0.83 0.85 0.80 0.62 0.85 0.70 0.36 5 threat_dir_h H 0.54 0.53 0.94 0.89 0.83 0.71 0.94 0.86 0.76 6 threat_norm_h H 0.52 0.74 0.98 0.93 0.84 0.78 0.94 0.97 0.87 7 slur_h H 1.00 0.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 8 profanity_h H 0.81 0.60 0.91 0.83 0.74 0.79 0.90 0.88 0.80 9 profanity_nh NH 1.00 0.86 0.57 0.71 0.86 0.71 0.71 0.86 0.86 10 ref_subs_clause_h H 0.65 0.58 0.96 0.92 0.75 0.70 0.90 0.73 0.56 11 ref_subs_sent_h H 0.85 0.83 1.00 0.99 0.98 0.97 1.00 0.99 0.91 12 negate_pos_h H 0.53 0.56 0.82 0.84 0.72 0.52 0.75 0.81 0.50 13 negate_neg_nh NH 0.87 0.76 0.44 0.68 0.72 0.70 0.63 0.87 0.90 14 phrase_question_h H 0.74 0.66 0.99 0.94 0.92 0.93 0.99 0.90 0.87 15 phrase_opinion_h H 0.91 0.86 0.98 0.98 0.97 0.96 0.96 0.99 0.96 16 ident_neutral_nh NH 0.93 1.00 0.81 0.91 0.97 0.93 0.75 0.96 1.00 17 ident_pos_nh NH 0.91 0.93 0.83 0.94 0.97 0.98 0.79 0.97 1.00 18 counter_quote_nh NH 0.18 0.32 0.02 0.04 0.08 0.11 0.08 0.32 0.29 19 counter_ref_nh NH 0.15 0.23 0.05 0.10 0.20 0.19 0.06 0.43 0.41 20 target_obj_nh NH 0.89 0.89 0.44 0.78 0.89 0.78 0.78 0.67 0.67 21 target_indiv_nh NH 0.56 0.67 0.00 0.11 0.33 0.44 0.22 0.22 0.33 22 target_group_nh NH 0.40 0.60 0.00 0.00 0.80 0.20 0.40 0.40 0.20 Table 35. Accuracy across fine-tuned models for different Functional Tests in Malay High-Quality Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia69 f_n t_function t_g Ministral Llama3b Llama8b sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.98 0.88 1.00 0.99 1.00 1.00 0.98 0.87 0.96 2 derog_neg_attrib_h H 1.00 0.93 1.00 1.00 0.98 1.00 1.00 0.99 0.99 3 derog_dehum_h H 1.00 0.87 1.00 1.00 0.98 0.98 0.98 0.97 0.98 4 derog_impl_h H 0.63 0.42 0.94 0.65 0.86 0.66 0.79 0.85 0.76 5 threat_dir_h H 0.17 0.38 0.79 0.67 0.92 0.79 0.71 0.75 0.71 6 threat_norm_h H 0.72 0.57 1.00 0.95 1.00 1.00 0.89 0.72 0.71 7 slur_h H 0.33 0.33 0.67 0.33 0.67 0.67 0.67 1.00 1.00 8 profanity_h H 1.00 0.85 1.00 1.00 1.00 1.00 1.00 0.99 0.99 9 profanity_nh NH 0.67 0.33 0.33 0.33 0.00 0.33 0.00 0.67 1.00 10 ref_subs_clause_h H 0.99 0.93 1.00 0.98 1.00 1.00 1.00 0.99 0.94 11 ref_subs_sent_h H 1.00 0.97 1.00 1.00 1.00 1.00 0.99 1.00 0.98 12 negate_pos_h H 0.79 0.94 1.00 0.98 0.95 0.86 0.82 0.93 0.92 13 negate_neg_nh NH 0.57 0.58 0.46 0.68 0.63 0.53 0.39 0.68 0.68 14 phrase_question_h H 0.98 0.94 1.00 0.99 1.00 1.00 1.00 1.00 1.00 15 phrase_opinion_h H 0.99 0.93 1.00 1.00 1.00 1.00 1.00 1.00 0.98 16 ident_neutral_nh NH 1.00 1.00 0.83 0.83 0.75 1.00 0.83 1.00 1.00 17 ident_pos_nh NH 0.97 0.95 0.95 0.95 0.97 0.97 0.97 0.97 0.97 18 counter_quote_nh NH 0.04 0.21 0.00 0.00 0.01 0.00 0.01 0.07 0.06 19 counter_ref_nh NH 0.13 0.03 0.00 0.01 0.02 0.05 0.03 0.06 0.02 20 target_obj_nh NH 0.44 0.55 0.63 0.61 0.71 0.59 0.70 0.60 0.59 21 target_indiv_nh NH 0.92 0.94 1.00 0.97 0.95 1.00 1.00 0.84 0.81 22 target_group_nh NH 0.45 0.44 0.84 0.71 0.84 0.82 0.88 0.75 0.71 Table 36. Accuracy across fine-tuned models for different Functional Tests in Mandarin High-Quality Test Cases. Manuscript submitted to ACM 70Ng et al. f_n t_function t_g Ministral Llama3b Llama8b sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.90 0.80 0.99 0.95 0.89 0.82 0.85 0.91 0.87 2 derog_neg_attrib_h H 0.86 0.70 0.99 0.97 0.93 0.95 0.93 0.93 0.89 3 derog_dehum_h H 0.82 0.68 0.97 0.92 0.87 0.86 0.88 0.96 0.84 4 derog_impl_h H 0.82 0.68 1.00 0.99 0.93 0.87 0.88 0.97 0.79 5 threat_dir_h H 0.89 0.81 0.95 0.92 0.92 0.93 0.98 0.96 0.93 6 threat_norm_h H 0.67 0.60 0.93 0.79 0.89 0.86 0.91 0.94 0.83 7 slur_h H 1.00 0.29 1.00 1.00 0.71 0.86 0.86 0.71 0.71 8 profanity_h H 0.93 0.81 1.00 0.94 0.96 0.94 0.95 0.99 0.89 9 profanity_nh NH 0.80 1.00 0.20 1.00 0.80 0.80 0.80 0.60 1.00 10 ref_subs_clause_h H 0.91 0.92 0.99 0.97 0.95 0.88 0.98 0.99 0.92 11 ref_subs_sent_h H 0.98 0.93 0.99 0.99 0.99 0.98 0.97 1.00 0.97 12 negate_pos_h H 0.80 0.66 0.93 0.93 0.79 0.72 0.77 0.95 0.81 13 negate_neg_nh NH 0.94 0.74 0.65 0.79 0.77 0.74 0.82 0.87 0.95 14 phrase_question_h H 0.92 0.93 1.00 0.99 0.97 0.95 0.95 0.98 0.98 15 phrase_opinion_h H 0.98 0.98 0.98 0.98 0.98 0.98 0.99 1.00 0.98 16 ident_neutral_nh NH 0.93 0.88 0.76 0.89 0.99 0.99 0.90 0.92 0.96 17 ident_pos_nh NH 0.96 0.99 0.93 0.97 0.97 1.00 0.96 0.99 1.00 18 counter_quote_nh NH 0.23 0.30 0.10 0.30 0.00 0.05 0.16 0.48 0.16 19 counter_ref_nh NH 0.48 0.25 0.12 0.42 0.22 0.23 0.34 0.40 0.32 20 target_obj_nh NH 0.78 0.78 1.00 0.97 0.86 0.83 0.93 1.00 0.91 21 target_indiv_nh NH 0.85 0.60 0.95 0.89 0.73 0.81 0.93 0.85 0.79 22 target_group_nh NH 0.88 0.79 0.98 0.98 0.82 0.84 0.88 0.95 0.75 Table 37. Accuracy across fine-tuned models for different Functional Tests in Singlish High-Quality Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia71 f_n t_function t_g Ministral Llama3b Llama8b sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.82 0.77 1.00 0.98 0.78 0.62 0.64 0.87 0.95 2 derog_neg_attrib_h H 0.88 0.75 0.98 0.90 0.77 0.79 0.83 0.91 0.94 3 derog_dehum_h H 0.83 0.55 1.00 0.95 0.66 0.72 0.91 0.93 0.96 4 derog_impl_h H 0.77 0.55 0.92 0.80 0.46 0.62 0.75 0.80 0.72 5 threat_dir_h H 0.73 0.56 0.91 0.79 0.45 0.80 0.91 0.88 0.91 6 threat_norm_h H 0.73 0.72 0.97 0.92 0.65 0.63 0.75 0.97 0.96 7 slur_h H 0.50 0.25 1.00 0.88 0.25 0.63 0.63 0.50 0.75 8 profanity_h H 0.59 0.53 0.99 0.95 0.58 0.71 0.67 0.89 0.87 9 profanity_nh NH 0.81 0.52 0.06 0.55 0.90 0.74 0.81 0.52 0.55 10 ref_subs_clause_h H 0.81 0.87 1.00 0.99 0.64 0.66 0.83 0.98 1.00 11 ref_subs_sent_h H 0.87 0.91 0.93 0.99 0.88 0.95 0.96 0.99 1.00 12 negate_pos_h H 0.64 0.66 0.99 0.95 0.78 0.71 0.90 0.93 0.95 13 negate_neg_nh NH 0.59 0.49 0.01 0.42 0.62 0.56 0.34 0.59 0.62 14 phrase_question_h H 0.89 0.91 1.00 0.99 0.76 0.80 0.94 0.97 0.98 15 phrase_opinion_h H 0.75 0.75 0.83 0.74 0.64 0.69 0.77 0.88 0.86 16 ident_neutral_nh NH 0.87 0.98 0.33 0.89 0.96 0.89 0.83 0.96 0.96 17 ident_pos_nh NH 0.93 1.00 0.80 1.00 0.94 0.90 0.70 0.98 0.99 18 counter_quote_nh NH 0.03 0.17 0.00 0.00 0.35 0.31 0.17 0.20 0.11 19 counter_ref_nh NH 0.11 0.34 0.00 0.07 0.23 0.20 0.36 0.16 0.00 20 target_obj_nh NH 0.48 0.61 0.00 0.35 0.55 0.42 0.58 0.45 0.55 21 target_indiv_nh NH 0.17 0.33 0.03 0.07 0.40 0.40 0.43 0.23 0.03 22 target_group_nh NH 0.35 0.39 0.04 0.22 0.52 0.39 0.30 0.17 0.17 Table 38. Accuracy across fine-tuned models for different Functional Tests in Tamil High-Quality Test Cases. Manuscript submitted to ACM 72Ng et al. f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.68 0.62 0.45 0.39 0.88 0.66 0.41 0.64 0.66 0.63 0.89 0.20 2 derog_neg_attrib_h H 0.75 0.65 0.44 0.38 0.93 0.65 0.47 0.76 0.81 0.80 0.97 0.37 3 derog_dehum_h H 0.87 0.77 0.62 0.62 0.96 0.79 0.72 0.87 0.85 0.86 0.95 0.62 4 derog_impl_h H 0.69 0.68 0.44 0.45 0.84 0.63 0.46 0.65 0.68 0.65 0.88 0.34 5 threat_dir_h H 0.86 0.87 0.66 0.71 0.94 0.82 0.66 0.87 0.87 0.74 0.87 0.54 6 threat_norm_h H 0.90 0.87 0.67 0.73 0.94 0.84 0.75 0.88 0.90 0.86 0.89 0.65 7 slur_h H 0.16 0.32 0.03 0.10 0.42 0.13 0.03 0.19 0.23 0.19 0.32 0.19 8 profanity_h H 0.77 0.69 0.49 0.48 0.92 0.75 0.53 0.72 0.77 0.73 0.94 0.42 10 ref_subs_clause_h H 0.83 0.77 0.54 0.63 0.89 0.79 0.62 0.77 0.79 0.74 0.93 0.52 11 ref_subs_sent_h H 0.85 0.79 0.60 0.64 0.97 0.86 0.52 0.85 0.89 0.84 0.96 0.50 12 negate_pos_h H 0.72 0.79 0.55 0.60 0.90 0.73 0.55 0.77 0.82 0.80 0.91 0.51 13 negate_neg_nh NH 0.77 0.80 0.84 0.92 0.75 0.77 0.92 0.88 0.88 0.92 0.87 0.99 14 phrase_question_h H 0.67 0.68 0.42 0.42 0.83 0.69 0.41 0.65 0.69 0.59 0.84 0.34 15 phrase_opinion_h H 0.72 0.66 0.42 0.52 0.86 0.68 0.56 0.75 0.77 0.75 0.86 0.50 16 ident_neutral_nh NH 0.88 0.81 0.95 0.99 0.73 0.83 0.95 0.95 0.95 0.97 0.89 0.99 17 ident_pos_nh NH 0.69 0.85 0.96 0.98 0.87 0.93 0.95 0.96 0.96 0.99 0.98 1.00 18 counter_quote_nh NH 0.49 0.55 0.56 0.68 0.43 0.45 0.77 0.76 0.64 0.71 0.58 0.89 19 counter_ref_nh NH 0.43 0.38 0.53 0.65 0.24 0.30 0.72 0.59 0.54 0.59 0.30 0.83 Table 39. Accuracy across non-finetuned models for different Functional Tests in Indonesian Silver Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia73 f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.80 0.67 0.69 0.52 0.93 0.60 0.51 0.62 0.71 0.52 0.87 0.32 2 derog_neg_attrib_h H 0.76 0.53 0.56 0.43 0.91 0.55 0.50 0.70 0.76 0.55 0.92 0.32 3 derog_dehum_h H 0.81 0.74 0.65 0.61 0.94 0.75 0.68 0.84 0.84 0.75 0.93 0.48 4 derog_impl_h H 0.71 0.57 0.50 0.49 0.93 0.53 0.56 0.59 0.67 0.47 0.81 0.23 5 threat_dir_h H 0.76 0.66 0.62 0.51 0.90 0.62 0.53 0.68 0.74 0.51 0.79 0.31 6 threat_norm_h H 0.97 0.86 0.80 0.78 0.99 0.80 0.72 0.93 0.95 0.80 0.94 0.69 7 slur_h H 0.35 0.26 0.28 0.11 0.61 0.26 0.13 0.20 0.20 0.13 0.24 0.07 8 profanity_h H 0.89 0.74 0.78 0.59 0.96 0.70 0.60 0.81 0.87 0.73 0.91 0.52 10 ref_subs_clause_h H 0.86 0.71 0.74 0.64 0.94 0.65 0.63 0.73 0.82 0.63 0.92 0.49 11 ref_subs_sent_h H 0.86 0.64 0.76 0.53 0.96 0.64 0.50 0.72 0.81 0.57 0.90 0.38 12 negate_pos_h H 0.72 0.52 0.60 0.54 0.88 0.62 0.53 0.63 0.68 0.51 0.82 0.28 13 negate_neg_nh NH 0.57 0.74 0.63 0.84 0.52 0.69 0.87 0.84 0.79 0.90 0.74 0.95 14 phrase_question_h H 0.78 0.60 0.66 0.41 0.89 0.61 0.29 0.63 0.70 0.46 0.87 0.29 15 phrase_opinion_h H 0.76 0.62 0.58 0.49 0.92 0.56 0.49 0.66 0.75 0.56 0.87 0.33 16 ident_neutral_nh NH 0.86 0.88 0.92 0.94 0.59 0.85 0.90 0.94 0.91 0.98 0.82 1.00 17 ident_pos_nh NH 0.73 0.82 0.89 0.97 0.71 0.84 0.91 0.95 0.93 0.97 0.91 1.00 18 counter_quote_nh NH 0.30 0.51 0.39 0.64 0.20 0.40 0.75 0.58 0.42 0.60 0.28 0.82 19 counter_ref_nh NH 0.32 0.63 0.49 0.68 0.20 0.47 0.71 0.61 0.53 0.67 0.33 0.87 Table 40. Accuracy across non-finetuned models for different Functional Tests in Tagalog Silver Test Cases. Manuscript submitted to ACM 74Ng et al. f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.69 0.57 0.65 0.44 0.68 0.42 0.68 0.68 0.70 0.68 0.83 0.40 2 derog_neg_attrib_h H 0.82 0.76 0.81 0.53 0.70 0.27 0.77 0.77 0.78 0.78 0.87 0.47 3 derog_dehum_h H 0.77 0.64 0.64 0.58 0.74 0.42 0.74 0.78 0.77 0.74 0.85 0.44 4 derog_impl_h H 0.72 0.63 0.60 0.51 0.72 0.35 0.70 0.67 0.74 0.74 0.86 0.38 5 threat_dir_h H 0.90 0.92 0.74 0.80 0.85 0.53 0.78 0.88 0.88 0.74 0.79 0.64 6 threat_norm_h H 0.94 0.90 0.88 0.85 0.80 0.61 0.89 0.92 0.93 0.90 0.90 0.66 7 slur_h H 0.32 0.28 0.34 0.21 0.44 0.17 0.33 0.35 0.38 0.43 0.55 0.12 8 profanity_h H 0.74 0.76 0.77 0.59 0.69 0.40 0.79 0.80 0.82 0.77 0.86 0.47 10 ref_subs_clause_h H 0.79 0.67 0.74 0.61 0.79 0.50 0.72 0.81 0.82 0.75 0.88 0.49 11 ref_subs_sent_h H 0.80 0.67 0.71 0.57 0.80 0.42 0.70 0.76 0.80 0.79 0.88 0.46 12 negate_pos_h H 0.85 0.79 0.73 0.71 0.79 0.49 0.82 0.83 0.86 0.84 0.88 0.53 13 negate_neg_nh NH 0.66 0.75 0.71 0.90 0.74 0.71 0.82 0.89 0.84 0.88 0.87 0.98 14 phrase_question_h H 0.74 0.73 0.69 0.52 0.68 0.35 0.69 0.72 0.75 0.73 0.84 0.37 15 phrase_opinion_h H 0.85 0.81 0.72 0.59 0.85 0.44 0.79 0.74 0.77 0.78 0.83 0.54 16 ident_neutral_nh NH 0.92 0.94 0.95 0.98 0.86 0.91 0.94 0.94 0.95 0.91 0.92 1.00 17 ident_pos_nh NH 0.83 0.90 0.93 0.98 0.85 0.87 0.93 0.97 0.97 0.99 0.98 1.00 18 counter_quote_nh NH 0.32 0.49 0.43 0.55 0.32 0.39 0.48 0.54 0.44 0.42 0.37 0.72 19 counter_ref_nh NH 0.46 0.67 0.47 0.68 0.47 0.58 0.64 0.68 0.61 0.64 0.58 0.91 Table 41. Accuracy across non-finetuned models for different Functional Tests in Thai Silver Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia75 f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.85 0.82 0.66 0.55 0.95 0.74 0.84 0.80 0.84 0.76 0.96 0.52 2 derog_neg_attrib_h H 0.91 0.84 0.73 0.66 0.98 0.77 0.91 0.87 0.89 0.85 0.95 0.61 3 derog_dehum_h H 0.88 0.85 0.76 0.75 0.97 0.78 0.89 0.87 0.89 0.89 0.95 0.71 4 derog_impl_h H 0.70 0.63 0.52 0.41 0.92 0.60 0.68 0.63 0.66 0.68 0.90 0.31 5 threat_dir_h H 0.93 0.94 0.81 0.75 0.98 0.83 0.75 0.89 0.89 0.80 0.85 0.64 6 threat_norm_h H 0.90 0.86 0.80 0.73 0.94 0.70 0.76 0.81 0.83 0.77 0.92 0.59 7 slur_h H 0.40 0.38 0.33 0.13 0.73 0.18 0.28 0.20 0.18 0.38 0.45 0.00 8 profanity_h H 0.75 0.66 0.60 0.53 0.96 0.58 0.73 0.72 0.74 0.72 0.92 0.45 10 ref_subs_clause_h H 0.90 0.83 0.78 0.62 0.96 0.81 0.84 0.86 0.88 0.81 0.93 0.61 11 ref_subs_sent_h H 0.89 0.75 0.70 0.64 0.98 0.71 0.82 0.78 0.82 0.77 0.94 0.52 12 negate_pos_h H 0.95 0.86 0.78 0.79 0.97 0.74 0.87 0.89 0.89 0.86 0.98 0.71 13 negate_neg_nh NH 0.84 0.89 0.89 0.98 0.80 0.84 0.96 0.99 0.98 0.95 0.98 0.99 14 phrase_question_h H 0.92 0.89 0.82 0.73 0.98 0.80 0.83 0.85 0.87 0.85 0.92 0.64 15 phrase_opinion_h H 0.90 0.87 0.78 0.69 0.98 0.69 0.86 0.85 0.86 0.84 0.97 0.64 16 ident_neutral_nh NH 0.94 0.92 0.96 1.00 0.78 0.89 0.92 0.95 0.96 0.96 0.88 1.00 17 ident_pos_nh NH 0.99 1.00 1.00 1.00 0.95 0.97 0.99 1.00 0.99 0.99 1.00 1.00 18 counter_quote_nh NH 0.33 0.56 0.46 0.60 0.21 0.36 0.51 0.59 0.46 0.42 0.38 0.75 19 counter_ref_nh NH 0.51 0.67 0.56 0.73 0.32 0.51 0.69 0.69 0.63 0.63 0.53 0.88 Table 42. Accuracy across non-finetuned models for different Functional Tests in Vietnamese Silver Test Cases. Manuscript submitted to ACM 76Ng et al. f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.82 0.77 0.68 0.45 0.89 0.77 0.59 0.76 0.75 0.85 0.90 0.41 2 derog_neg_attrib_h H 0.78 0.73 0.61 0.36 0.90 0.75 0.59 0.81 0.82 0.88 0.92 0.42 3 derog_dehum_h H 0.88 0.82 0.81 0.59 0.97 0.84 0.73 0.86 0.89 0.94 0.94 0.61 4 derog_impl_h H 0.81 0.72 0.68 0.60 0.87 0.78 0.67 0.83 0.82 0.81 0.92 0.53 5 threat_dir_h H 0.93 0.90 0.85 0.77 0.96 0.86 0.82 0.86 0.90 0.90 0.97 0.77 6 threat_norm_h H 0.93 0.88 0.87 0.77 0.88 0.83 0.85 0.91 0.91 0.94 0.94 0.79 7 slur_h H 0.51 0.47 0.33 0.09 0.76 0.58 0.36 0.49 0.47 0.51 0.62 0.16 8 profanity_h H 0.85 0.78 0.67 0.45 0.90 0.80 0.60 0.80 0.80 0.85 0.93 0.50 10 ref_subs_clause_h H 0.92 0.88 0.86 0.70 0.99 0.88 0.79 0.90 0.91 0.91 0.97 0.72 11 ref_subs_sent_h H 0.84 0.75 0.75 0.57 0.98 0.85 0.74 0.86 0.85 0.86 0.91 0.57 12 negate_pos_h H 0.90 0.77 0.68 0.51 0.90 0.77 0.71 0.81 0.82 0.86 0.93 0.54 13 negate_neg_nh NH 0.55 0.58 0.62 0.92 0.51 0.54 0.81 0.75 0.80 0.86 0.82 0.96 14 phrase_question_h H 0.90 0.85 0.79 0.61 0.94 0.86 0.73 0.80 0.81 0.82 0.88 0.59 15 phrase_opinion_h H 0.82 0.77 0.72 0.52 0.92 0.81 0.65 0.81 0.84 0.88 0.94 0.60 16 ident_neutral_nh NH 0.76 0.76 0.82 0.95 0.56 0.69 0.78 0.78 0.77 0.81 0.76 0.95 17 ident_pos_nh NH 0.82 0.86 0.90 0.98 0.75 0.81 0.93 0.93 0.95 0.96 0.93 1.00 18 counter_quote_nh NH 0.17 0.26 0.27 0.49 0.16 0.18 0.39 0.26 0.24 0.26 0.24 0.55 19 counter_ref_nh NH 0.24 0.36 0.38 0.59 0.13 0.22 0.52 0.38 0.35 0.33 0.27 0.70 Table 43. Accuracy across non-finetuned models for different Functional Tests in Malay Silver Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia77 f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.79 0.77 0.77 0.54 0.89 0.72 0.72 0.62 0.67 0.80 0.93 0.42 2 derog_neg_attrib_h H 0.68 0.81 0.78 0.53 0.82 0.77 0.66 0.59 0.63 0.74 0.77 0.42 3 derog_dehum_h H 0.84 0.79 0.85 0.65 0.94 0.77 0.81 0.76 0.78 0.89 0.87 0.58 4 derog_impl_h H 0.76 0.68 0.68 0.31 0.79 0.69 0.71 0.51 0.59 0.86 0.95 0.35 5 threat_dir_h H 0.84 0.81 0.87 0.54 0.96 0.79 0.77 0.78 0.79 0.90 0.87 0.47 6 threat_norm_h H 0.82 0.78 0.82 0.62 0.92 0.78 0.79 0.68 0.70 0.84 0.96 0.50 7 slur_h H 0.29 0.17 0.13 0.04 0.42 0.21 0.21 0.13 0.13 0.25 0.29 0.00 8 profanity_h H 0.94 0.93 0.94 0.78 0.98 0.86 0.94 0.88 0.89 0.94 0.97 0.71 10 ref_subs_clause_h H 0.89 0.85 0.85 0.68 0.96 0.86 0.83 0.77 0.85 0.87 0.92 0.57 11 ref_subs_sent_h H 0.76 0.72 0.75 0.51 0.82 0.67 0.60 0.60 0.65 0.71 0.69 0.42 12 negate_pos_h H 0.85 0.83 0.84 0.56 0.89 0.72 0.81 0.72 0.75 0.86 0.86 0.40 13 negate_neg_nh NH 0.70 0.73 0.61 0.96 0.84 0.71 0.80 0.93 0.92 0.94 0.97 0.99 14 phrase_question_h H 0.83 0.81 0.83 0.67 0.86 0.83 0.74 0.74 0.77 0.83 0.79 0.52 15 phrase_opinion_h H 0.87 0.92 0.93 0.77 0.94 0.88 0.86 0.86 0.87 0.91 0.93 0.69 16 ident_neutral_nh NH 0.96 0.89 0.92 0.97 0.74 0.85 0.93 0.95 0.94 0.93 0.84 1.00 17 ident_pos_nh NH 0.94 0.97 0.94 0.98 0.93 0.94 0.95 0.97 0.97 0.98 0.96 0.99 18 counter_quote_nh NH 0.38 0.53 0.38 0.67 0.39 0.47 0.49 0.67 0.52 0.38 0.48 0.78 19 counter_ref_nh NH 0.37 0.40 0.32 0.64 0.39 0.37 0.63 0.62 0.55 0.48 0.55 0.77 Table 44. Accuracy across non-finetuned models for different Functional Tests in Mandarin Silver Test Cases. Manuscript submitted to ACM 78Ng et al. f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.78 0.94 0.72 0.46 0.85 0.69 0.74 0.71 0.73 0.87 0.92 0.51 2 derog_neg_attrib_h H 0.82 0.98 0.81 0.57 0.91 0.75 0.79 0.75 0.82 0.87 0.89 0.65 3 derog_dehum_h H 0.93 0.99 0.97 0.78 0.93 0.83 0.91 0.92 0.95 0.97 0.97 0.79 4 derog_impl_h H 0.80 0.85 0.83 0.62 0.80 0.71 0.80 0.73 0.75 0.86 0.96 0.68 5 threat_dir_h H 0.74 0.92 0.75 0.54 0.85 0.69 0.77 0.72 0.74 0.84 0.97 0.61 6 threat_norm_h H 0.74 0.85 0.72 0.42 0.88 0.70 0.76 0.68 0.74 0.90 0.95 0.53 7 slur_h H 0.45 0.59 0.59 0.27 0.82 0.64 0.41 0.45 0.55 0.68 0.59 0.59 8 profanity_h H 0.75 0.87 0.66 0.46 0.81 0.65 0.70 0.72 0.73 0.86 0.89 0.54 10 ref_subs_clause_h H 0.84 0.95 0.86 0.67 0.66 0.48 0.87 0.81 0.86 0.96 0.96 0.77 11 ref_subs_sent_h H 0.70 0.81 0.74 0.48 0.75 0.60 0.70 0.70 0.71 0.83 0.86 0.56 12 negate_pos_h H 0.84 0.89 0.84 0.68 0.94 0.82 0.84 0.83 0.86 0.95 0.96 0.70 13 negate_neg_nh NH 0.83 0.72 0.74 0.98 0.64 0.57 0.88 0.88 0.91 0.93 0.86 0.98 14 phrase_question_h H 0.86 0.95 0.91 0.66 0.74 0.61 0.88 0.82 0.85 0.93 0.95 0.71 15 phrase_opinion_h H 0.83 0.93 0.83 0.64 0.90 0.77 0.87 0.79 0.84 0.96 0.97 0.73 16 ident_neutral_nh NH 0.77 0.71 0.79 0.91 0.50 0.61 0.81 0.78 0.80 0.81 0.68 0.92 17 ident_pos_nh NH 0.89 0.86 0.87 0.98 0.70 0.70 0.88 0.94 0.93 0.96 0.91 1.00 18 counter_quote_nh NH 0.28 0.20 0.18 0.54 0.21 0.21 0.36 0.36 0.30 0.34 0.19 0.51 19 counter_ref_nh NH 0.36 0.26 0.26 0.62 0.29 0.19 0.45 0.51 0.41 0.49 0.44 0.69 Table 45. Accuracy across non-finetuned models for different Functional Tests in Singlish Silver Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia79 f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek 1 derog_neg_emote_h H 0.84 0.73 0.99 0.70 0.95 0.56 0.49 0.76 0.90 0.95 0.72 0.45 2 derog_neg_attrib_h H 0.83 0.81 0.91 0.79 0.99 0.60 0.68 0.81 0.89 0.95 0.84 0.59 3 derog_dehum_h H 0.94 0.90 0.99 0.91 1.00 0.68 0.74 0.90 0.96 0.97 0.95 0.71 4 derog_impl_h H 0.61 0.57 0.81 0.47 0.86 0.47 0.45 0.56 0.59 0.70 0.65 0.19 5 threat_dir_h H 0.86 0.74 0.94 0.80 0.91 0.54 0.58 0.82 0.90 0.93 0.81 0.67 6 threat_norm_h H 0.92 0.85 0.97 0.87 0.99 0.77 0.69 0.92 0.96 0.99 0.95 0.78 7 slur_h H 0.17 0.33 0.93 0.43 0.87 0.07 0.20 0.17 0.40 0.57 0.27 0.03 8 profanity_h H 0.83 0.68 0.97 0.73 0.97 0.58 0.58 0.75 0.83 0.93 0.88 0.46 10 ref_subs_clause_h H 0.76 0.62 0.93 0.67 0.95 0.54 0.55 0.69 0.76 0.77 0.76 0.38 11 ref_subs_sent_h H 0.86 0.73 0.99 0.81 0.99 0.67 0.57 0.74 0.86 0.90 0.87 0.63 12 negate_pos_h H 0.82 0.75 0.96 0.80 1.00 0.60 0.72 0.86 0.92 0.92 0.88 0.51 13 negate_neg_nh NH 0.41 0.54 0.10 0.59 0.33 0.51 0.58 0.66 0.62 0.51 0.95 0.93 14 phrase_question_h H 0.74 0.72 0.95 0.66 0.95 0.61 0.38 0.74 0.79 0.85 0.82 0.42 15 phrase_opinion_h H 0.93 0.85 0.99 0.82 0.99 0.59 0.67 0.88 0.92 0.96 0.94 0.60 16 ident_neutral_nh NH 0.63 0.70 0.38 0.90 0.23 0.59 0.65 0.96 0.96 0.82 0.96 1.00 17 ident_pos_nh NH 0.52 0.74 0.42 0.89 0.25 0.61 0.51 0.81 0.86 0.88 0.96 1.00 18 counter_quote_nh NH 0.23 0.44 0.09 0.41 0.06 0.30 0.61 0.51 0.34 0.31 0.59 0.72 19 counter_ref_nh NH 0.28 0.40 0.07 0.52 0.08 0.33 0.63 0.50 0.43 0.48 0.56 0.81 Table 46. Accuracy across non-finetuned models for different Functional Tests in Tamil Silver Test Cases. Manuscript submitted to ACM 80Ng et al. f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.40 0.39 0.40 0.57 0.51 0.39 0.52 0.56 0.38 2 derog_neg_attrib_h H 0.45 0.39 0.47 0.66 0.55 0.44 0.61 0.63 0.49 3 derog_dehum_h H 0.58 0.54 0.62 0.76 0.70 0.62 0.77 0.73 0.65 4 derog_impl_h H 0.38 0.32 0.37 0.61 0.46 0.38 0.50 0.51 0.41 5 threat_dir_h H 0.48 0.53 0.56 0.62 0.65 0.56 0.53 0.64 0.52 6 threat_norm_h H 0.52 0.63 0.65 0.75 0.71 0.69 0.71 0.76 0.65 7 slur_h H 0.16 0.23 0.23 0.23 0.19 0.16 0.16 0.19 0.23 8 profanity_h H 0.53 0.49 0.52 0.68 0.60 0.53 0.64 0.63 0.52 10 ref_subs_clause_h H 0.48 0.45 0.48 0.67 0.65 0.51 0.64 0.64 0.48 11 ref_subs_sent_h H 0.59 0.53 0.57 0.79 0.68 0.61 0.75 0.72 0.56 12 negate_pos_h H 0.53 0.47 0.53 0.70 0.64 0.45 0.65 0.72 0.60 13 negate_neg_nh NH 0.93 0.93 0.94 0.94 0.94 0.94 0.93 0.94 0.96 14 phrase_question_h H 0.41 0.35 0.38 0.54 0.46 0.38 0.47 0.56 0.38 15 phrase_opinion_h H 0.43 0.39 0.46 0.60 0.57 0.45 0.57 0.61 0.51 16 ident_neutral_nh NH 0.97 0.97 0.98 0.98 0.98 0.98 0.97 0.98 0.99 17 ident_pos_nh NH 0.90 0.98 0.97 0.99 0.99 0.99 0.99 0.98 1.00 18 counter_quote_nh NH 0.71 0.76 0.85 0.76 0.73 0.81 0.74 0.81 0.89 19 counter_ref_nh NH 0.62 0.65 0.73 0.61 0.65 0.79 0.64 0.66 0.78 Table 47. Accuracy across fine-tuned models for different Functional Tests in Indonesian Silver Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia81 f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.89 0.80 0.95 0.91 0.90 0.85 0.89 0.89 0.85 2 derog_neg_attrib_h H 0.90 0.80 0.92 0.92 0.85 0.76 0.84 0.92 0.85 3 derog_dehum_h H 0.86 0.87 0.96 0.94 0.91 0.88 0.91 0.90 0.87 4 derog_impl_h H 0.72 0.71 0.92 0.88 0.84 0.75 0.86 0.85 0.77 5 threat_dir_h H 0.71 0.77 0.81 0.77 0.78 0.71 0.79 0.80 0.75 6 threat_norm_h H 0.86 0.88 0.97 0.97 0.97 0.85 0.86 0.96 0.94 7 slur_h H 0.33 0.28 0.54 0.48 0.35 0.37 0.52 0.39 0.37 8 profanity_h H 0.94 0.93 0.97 0.96 0.95 0.93 0.97 0.94 0.93 10 ref_subs_clause_h H 0.89 0.86 0.94 0.93 0.91 0.86 0.90 0.91 0.89 11 ref_subs_sent_h H 0.92 0.88 0.95 0.96 0.89 0.86 0.91 0.92 0.89 12 negate_pos_h H 0.78 0.72 0.86 0.85 0.78 0.72 0.80 0.84 0.76 13 negate_neg_nh NH 0.71 0.73 0.62 0.74 0.71 0.76 0.67 0.72 0.78 14 phrase_question_h H 0.78 0.72 0.91 0.86 0.80 0.76 0.76 0.86 0.82 15 phrase_opinion_h H 0.84 0.81 0.91 0.89 0.85 0.80 0.83 0.88 0.80 16 ident_neutral_nh NH 0.83 0.89 0.73 0.84 0.87 0.87 0.78 0.88 0.91 17 ident_pos_nh NH 0.83 0.89 0.87 0.91 0.91 0.87 0.89 0.91 0.92 18 counter_quote_nh NH 0.30 0.36 0.16 0.24 0.25 0.38 0.27 0.25 0.31 19 counter_ref_nh NH 0.38 0.42 0.24 0.34 0.36 0.45 0.31 0.31 0.43 Table 48. Accuracy across fine-tuned models for different Functional Tests in Tagalog Silver Test Cases. Manuscript submitted to ACM 82Ng et al. f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.47 0.55 0.75 0.72 0.74 0.59 0.67 0.55 0.49 2 derog_neg_attrib_h H 0.44 0.68 0.85 0.85 0.78 0.53 0.77 0.79 0.76 3 derog_dehum_h H 0.60 0.54 0.76 0.71 0.70 0.57 0.68 0.70 0.62 4 derog_impl_h H 0.47 0.50 0.73 0.69 0.72 0.51 0.62 0.67 0.57 5 threat_dir_h H 0.57 0.66 0.74 0.76 0.73 0.63 0.62 0.69 0.62 6 threat_norm_h H 0.81 0.78 0.90 0.90 0.88 0.77 0.87 0.89 0.79 7 slur_h H 0.13 0.21 0.35 0.36 0.26 0.14 0.22 0.26 0.22 8 profanity_h H 0.43 0.69 0.86 0.87 0.76 0.58 0.72 0.59 0.72 10 ref_subs_clause_h H 0.68 0.69 0.82 0.82 0.84 0.68 0.70 0.70 0.63 11 ref_subs_sent_h H 0.58 0.69 0.81 0.80 0.78 0.64 0.70 0.75 0.65 12 negate_pos_h H 0.54 0.69 0.84 0.82 0.80 0.57 0.75 0.79 0.70 13 negate_neg_nh NH 0.95 0.88 0.85 0.89 0.87 0.92 0.89 0.91 0.93 14 phrase_question_h H 0.51 0.55 0.73 0.75 0.71 0.56 0.66 0.68 0.64 15 phrase_opinion_h H 0.51 0.71 0.81 0.79 0.81 0.62 0.74 0.74 0.73 16 ident_neutral_nh NH 0.99 0.99 0.99 0.99 0.99 0.98 0.99 0.98 0.99 17 ident_pos_nh NH 1.00 1.00 1.00 1.00 0.99 1.00 1.00 1.00 1.00 18 counter_quote_nh NH 0.64 0.56 0.44 0.47 0.44 0.55 0.56 0.67 0.70 19 counter_ref_nh NH 0.87 0.71 0.63 0.71 0.60 0.77 0.70 0.77 0.79 Table 49. Accuracy across fine-tuned models for different Functional Tests in Thai Silver Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia83 f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.94 0.96 0.99 0.98 0.98 0.94 0.98 0.94 0.96 2 derog_neg_attrib_h H 0.96 0.93 0.98 0.98 0.98 0.96 0.96 0.98 0.97 3 derog_dehum_h H 0.96 0.97 0.98 0.99 0.96 0.96 0.96 0.96 0.96 4 derog_impl_h H 0.79 0.83 0.87 0.89 0.88 0.82 0.87 0.80 0.77 5 threat_dir_h H 0.92 0.88 0.95 0.93 0.93 0.91 0.95 0.90 0.92 6 threat_norm_h H 0.91 0.91 0.94 0.94 0.96 0.93 0.95 0.92 0.93 7 slur_h H 0.70 0.68 0.75 0.75 0.70 0.53 0.70 0.60 0.55 8 profanity_h H 0.92 0.91 0.91 0.94 0.93 0.90 0.96 0.89 0.91 10 ref_subs_clause_h H 0.95 0.96 0.96 0.96 0.97 0.94 0.97 0.95 0.97 11 ref_subs_sent_h H 0.94 0.94 0.94 0.95 0.97 0.92 0.96 0.96 0.94 12 negate_pos_h H 0.98 0.98 0.98 0.99 0.99 0.96 0.99 0.96 0.97 13 negate_neg_nh NH 0.90 0.92 0.91 0.95 0.92 0.94 0.85 0.97 0.98 14 phrase_question_h H 0.95 0.94 0.96 0.96 0.98 0.95 0.96 0.94 0.94 15 phrase_opinion_h H 0.98 0.96 0.98 0.99 0.98 0.95 0.97 0.98 0.97 16 ident_neutral_nh NH 0.92 0.91 0.86 0.88 0.91 0.96 0.91 0.92 0.93 17 ident_pos_nh NH 0.99 1.00 1.00 1.00 1.00 1.00 0.99 1.00 1.00 18 counter_quote_nh NH 0.29 0.29 0.29 0.31 0.27 0.30 0.20 0.46 0.41 19 counter_ref_nh NH 0.47 0.44 0.47 0.49 0.44 0.49 0.39 0.58 0.60 Table 50. Accuracy across fine-tuned models for different Functional Tests in Vietnamese Silver Test Cases. Manuscript submitted to ACM 84Ng et al. f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.67 0.50 0.91 0.85 0.74 0.66 0.85 0.77 0.61 2 derog_neg_attrib_h H 0.66 0.45 0.90 0.81 0.61 0.64 0.86 0.79 0.61 3 derog_dehum_h H 0.73 0.62 0.95 0.92 0.80 0.73 0.90 0.83 0.74 4 derog_impl_h H 0.75 0.62 0.92 0.86 0.80 0.70 0.89 0.72 0.60 5 threat_dir_h H 0.79 0.76 0.94 0.84 0.87 0.78 0.90 0.85 0.71 6 threat_norm_h H 0.64 0.67 0.94 0.85 0.85 0.80 0.91 0.87 0.79 7 slur_h H 0.33 0.33 0.67 0.56 0.33 0.36 0.56 0.44 0.20 8 profanity_h H 0.68 0.51 0.90 0.85 0.77 0.72 0.89 0.78 0.70 10 ref_subs_clause_h H 0.82 0.70 0.95 0.92 0.87 0.79 0.94 0.88 0.80 11 ref_subs_sent_h H 0.71 0.60 0.95 0.94 0.83 0.81 0.93 0.80 0.72 12 negate_pos_h H 0.75 0.60 0.91 0.84 0.70 0.70 0.86 0.75 0.66 13 negate_neg_nh NH 0.85 0.89 0.68 0.84 0.81 0.86 0.76 0.91 0.98 14 phrase_question_h H 0.74 0.63 0.89 0.85 0.81 0.79 0.88 0.76 0.71 15 phrase_opinion_h H 0.72 0.59 0.96 0.87 0.81 0.75 0.91 0.86 0.73 16 ident_neutral_nh NH 0.88 0.97 0.70 0.85 0.89 0.89 0.83 0.84 0.93 17 ident_pos_nh NH 0.95 0.98 0.92 0.98 0.98 0.99 0.93 0.97 0.99 18 counter_quote_nh NH 0.34 0.51 0.11 0.19 0.26 0.40 0.23 0.36 0.52 19 counter_ref_nh NH 0.48 0.54 0.18 0.29 0.38 0.46 0.28 0.41 0.52 Table 51. Accuracy across fine-tuned models for different Functional Tests in Malay Silver Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia85 f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.77 0.79 0.93 0.89 0.80 0.77 0.80 0.75 0.72 2 derog_neg_attrib_h H 0.76 0.69 0.81 0.79 0.79 0.77 0.77 0.72 0.69 3 derog_dehum_h H 0.86 0.80 0.96 0.92 0.89 0.83 0.89 0.86 0.77 4 derog_impl_h H 0.56 0.55 0.77 0.67 0.77 0.59 0.65 0.65 0.47 5 threat_dir_h H 0.86 0.79 0.96 0.91 0.91 0.88 0.87 0.91 0.80 6 threat_norm_h H 0.74 0.72 0.88 0.84 0.87 0.78 0.78 0.78 0.62 7 slur_h H 0.21 0.17 0.29 0.29 0.13 0.13 0.21 0.21 0.17 8 profanity_h H 0.90 0.84 0.94 0.93 0.95 0.89 0.91 0.91 0.88 10 ref_subs_clause_h H 0.87 0.89 0.97 0.96 0.90 0.90 0.93 0.92 0.89 11 ref_subs_sent_h H 0.79 0.78 0.91 0.90 0.80 0.78 0.79 0.78 0.71 12 negate_pos_h H 0.76 0.77 0.91 0.89 0.85 0.80 0.79 0.83 0.68 13 negate_neg_nh NH 0.89 0.89 0.79 0.92 0.90 0.88 0.74 0.94 0.92 14 phrase_question_h H 0.79 0.79 0.87 0.87 0.85 0.81 0.84 0.79 0.77 15 phrase_opinion_h H 0.86 0.82 0.95 0.93 0.94 0.90 0.88 0.90 0.86 16 ident_neutral_nh NH 0.91 0.94 0.77 0.82 0.90 0.92 0.86 0.87 0.92 17 ident_pos_nh NH 0.97 0.99 0.95 0.97 0.98 0.99 0.98 0.98 0.99 18 counter_quote_nh NH 0.33 0.39 0.23 0.28 0.25 0.25 0.20 0.43 0.44 19 counter_ref_nh NH 0.44 0.40 0.28 0.37 0.31 0.36 0.32 0.43 0.44 Table 52. Accuracy across fine-tuned models for different Functional Tests in Mandarin Silver Test Cases. Manuscript submitted to ACM 86Ng et al. f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.78 0.69 0.94 0.92 0.86 0.81 0.83 0.94 0.90 2 derog_neg_attrib_h H 0.86 0.76 0.97 0.96 0.95 0.89 0.90 0.94 0.91 3 derog_dehum_h H 0.92 0.87 0.99 0.95 0.97 0.91 0.97 0.99 0.95 4 derog_impl_h H 0.79 0.78 0.94 0.91 0.88 0.82 0.83 0.89 0.83 5 threat_dir_h H 0.83 0.77 0.97 0.93 0.92 0.81 0.83 0.93 0.86 6 threat_norm_h H 0.70 0.59 0.90 0.83 0.86 0.70 0.77 0.92 0.83 7 slur_h H 0.68 0.32 0.91 0.73 0.73 0.45 0.41 0.77 0.55 8 profanity_h H 0.83 0.65 0.97 0.93 0.85 0.79 0.80 0.94 0.89 10 ref_subs_clause_h H 0.83 0.82 0.97 0.96 0.94 0.82 0.92 0.95 0.90 11 ref_subs_sent_h H 0.75 0.72 0.93 0.91 0.87 0.77 0.79 0.92 0.86 12 negate_pos_h H 0.84 0.78 0.95 0.93 0.87 0.82 0.84 0.94 0.86 13 negate_neg_nh NH 0.93 0.92 0.83 0.88 0.83 0.88 0.90 0.87 0.92 14 phrase_question_h H 0.86 0.84 0.98 0.93 0.92 0.86 0.94 0.96 0.93 15 phrase_opinion_h H 0.87 0.72 0.96 0.94 0.90 0.87 0.90 0.94 0.92 16 ident_neutral_nh NH 0.86 0.92 0.69 0.81 0.79 0.88 0.83 0.78 0.90 17 ident_pos_nh NH 0.94 0.99 0.91 0.95 0.94 0.96 0.97 0.91 0.97 18 counter_quote_nh NH 0.32 0.41 0.17 0.25 0.19 0.23 0.31 0.29 0.35 19 counter_ref_nh NH 0.48 0.59 0.36 0.46 0.28 0.39 0.50 0.44 0.50 Table 53. Accuracy across fine-tuned models for different Functional Tests in Singlish Silver Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia87 f_n t_function t_g Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem 1 derog_neg_emote_h H 0.82 0.80 0.98 0.97 0.78 0.62 0.74 0.91 0.93 2 derog_neg_attrib_h H 0.86 0.75 0.96 0.89 0.67 0.72 0.76 0.89 0.92 3 derog_dehum_h H 0.85 0.64 0.99 0.94 0.69 0.78 0.83 0.92 0.92 4 derog_impl_h H 0.55 0.46 0.92 0.70 0.42 0.44 0.66 0.66 0.56 5 threat_dir_h H 0.75 0.66 0.91 0.82 0.54 0.73 0.86 0.87 0.88 6 threat_norm_h H 0.86 0.80 0.97 0.95 0.75 0.74 0.71 0.97 0.94 7 slur_h H 0.40 0.33 1.00 0.83 0.47 0.50 0.63 0.53 0.60 8 profanity_h H 0.59 0.44 0.97 0.83 0.53 0.65 0.59 0.84 0.78 10 ref_subs_clause_h H 0.70 0.62 0.93 0.85 0.50 0.48 0.74 0.85 0.80 11 ref_subs_sent_h H 0.82 0.84 0.95 0.95 0.70 0.84 0.86 0.93 0.88 12 negate_pos_h H 0.72 0.68 0.96 0.94 0.71 0.68 0.81 0.88 0.91 13 negate_neg_nh NH 0.62 0.57 0.12 0.54 0.68 0.61 0.42 0.71 0.76 14 phrase_question_h H 0.74 0.74 0.95 0.88 0.66 0.72 0.81 0.86 0.83 15 phrase_opinion_h H 0.80 0.79 0.94 0.90 0.65 0.71 0.76 0.90 0.90 16 ident_neutral_nh NH 0.87 0.98 0.33 0.94 0.96 0.87 0.87 0.96 0.99 17 ident_pos_nh NH 0.93 0.97 0.85 0.98 0.95 0.92 0.72 0.98 0.99 18 counter_quote_nh NH 0.33 0.45 0.08 0.24 0.39 0.38 0.34 0.41 0.34 19 counter_ref_nh NH 0.46 0.48 0.06 0.28 0.46 0.44 0.29 0.46 0.40 Table 54. Accuracy across fine-tuned models for different Functional Tests in Tamil Silver Test Cases. Manuscript submitted to ACM 88Ng et al. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Ethnicity/Race/Origin 0.66 0.67 0.63 0.70 0.71 0.68 0.73 0.80 0.78 0.85 0.89 0.73 Gender/Sexuality 0.50 0.53 0.55 0.65 0.52 0.44 0.63 0.69 0.69 0.67 0.77 0.61 Religion 0.66 0.64 0.69 0.79 0.70 0.67 0.74 0.79 0.78 0.84 0.89 0.78 Table 55. F1 Scores across non-finetuned models for different protected categories in Indonesian High-Quality Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Disability 0.61 0.53 0.49 0.41 0.52 0.43 0.53 0.74 0.72 0.64 0.79 0.47 Ethnicity/Race/Origin 0.66 0.64 0.60 0.59 0.56 0.51 0.60 0.77 0.77 0.81 0.86 0.68 Gender/Sexuality 0.67 0.68 0.63 0.70 0.61 0.53 0.67 0.79 0.81 0.82 0.85 0.74 PLHIV 0.58 0.53 0.67 0.78 0.53 0.57 0.77 0.72 0.70 0.72 0.84 0.82 Religion 0.68 0.66 0.63 0.66 0.59 0.54 0.70 0.79 0.79 0.81 0.80 0.68 Table 56. F1 Scores across non-finetuned models for different protected categories in Tagalog High-Quality Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Age 0.65 0.71 0.70 0.71 0.66 0.28 0.69 0.71 0.73 0.71 0.89 0.51 Disability 0.68 0.68 0.71 0.76 0.76 0.69 0.71 0.80 0.77 0.77 0.86 0.83 Ethnicity/Race/Origin 0.67 0.73 0.70 0.68 0.71 0.56 0.74 0.76 0.74 0.75 0.78 0.67 Gender/Sexuality 0.70 0.72 0.67 0.72 0.69 0.53 0.70 0.76 0.75 0.76 0.86 0.60 Religion 0.63 0.70 0.64 0.70 0.67 0.61 0.71 0.74 0.69 0.78 0.88 0.75 Vulnerable Workers 0.70 0.71 0.64 0.65 0.70 0.53 0.72 0.79 0.77 0.74 0.53 0.62 Table 57. F1 Scores across non-finetuned models for different protected categories in Thai High-Quality Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Age 0.74 0.70 0.61 0.56 0.80 0.57 0.73 0.80 0.80 0.69 0.89 0.53 Disability 0.76 0.72 0.69 0.72 0.77 0.80 0.73 0.85 0.84 0.78 0.85 0.70 Ethnicity/Race/Origin 0.77 0.76 0.74 0.76 0.75 0.81 0.81 0.84 0.82 0.84 0.87 0.85 Gender/Sexuality 0.75 0.73 0.68 0.73 0.73 0.74 0.75 0.81 0.81 0.77 0.87 0.68 PLHIV 0.68 0.70 0.76 0.84 0.76 0.85 0.75 0.80 0.81 0.69 0.90 0.89 Religion 0.74 0.77 0.80 0.81 0.69 0.62 0.77 0.84 0.82 0.84 0.89 0.86 Table 58. F1 Scores across non-finetuned models for different protected categories in Vietnamese High-Quality Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia89 p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Age 0.56 0.64 0.57 0.43 0.70 0.52 0.58 0.71 0.71 0.83 0.86 0.40 Disability 0.64 0.60 0.59 0.54 0.69 0.68 0.64 0.76 0.76 0.78 0.83 0.72 Ethnicity/Race/Origin 0.63 0.65 0.67 0.80 0.59 0.65 0.73 0.77 0.74 0.76 0.84 0.84 Gender/Sexuality 0.66 0.57 0.65 0.72 0.63 0.59 0.65 0.79 0.80 0.80 0.85 0.77 Religion 0.69 0.67 0.69 0.77 0.67 0.70 0.70 0.79 0.79 0.82 0.88 0.82 Table 59. F1 Scores across non-finetuned models for different protected categories in Malay High-Quality Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Age 0.57 0.64 0.60 0.43 0.74 0.38 0.62 0.62 0.63 0.75 0.85 0.51 Disability 0.66 0.64 0.63 0.71 0.73 0.79 0.72 0.83 0.75 0.73 0.85 0.70 Ethnicity/Race/Origin 0.62 0.68 0.57 0.73 0.64 0.72 0.75 0.77 0.72 0.75 0.83 0.71 Gender/Sexuality 0.71 0.67 0.61 0.67 0.65 0.66 0.69 0.80 0.75 0.79 0.84 0.69 Religion 0.61 0.64 0.62 0.69 0.62 0.67 0.73 0.77 0.69 0.69 0.80 0.65 Table 60. F1 Scores across non-finetuned models for different protected categories in Mandarin High-Quality Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Age 0.53 0.75 0.73 0.45 0.68 0.36 0.60 0.58 0.63 0.79 0.87 0.48 Disability 0.70 0.67 0.72 0.67 0.68 0.65 0.70 0.65 0.68 0.69 0.70 0.82 Ethnicity/Race/Origin 0.76 0.69 0.75 0.75 0.67 0.69 0.77 0.79 0.79 0.80 0.84 0.83 Gender/Sexuality 0.69 0.69 0.68 0.71 0.65 0.58 0.68 0.71 0.74 0.77 0.80 0.74 Religion 0.77 0.68 0.75 0.80 0.69 0.67 0.79 0.83 0.83 0.84 0.87 0.81 Table 61. F1 Scores across non-finetuned models for different protected categories in Singlish High-Quality Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Age 0.54 0.61 0.70 0.44 0.66 0.31 0.49 0.52 0.67 0.78 0.84 0.32 Disability 0.57 0.61 0.60 0.63 0.57 0.46 0.47 0.72 0.81 0.79 0.89 0.61 Ethnicity/Race/Origin 0.62 0.63 0.59 0.76 0.51 0.54 0.54 0.78 0.81 0.77 0.85 0.64 Gender/Sexuality 0.61 0.64 0.52 0.74 0.57 0.54 0.54 0.77 0.80 0.82 0.75 0.66 Religion 0.65 0.68 0.57 0.75 0.51 0.53 0.53 0.82 0.80 0.80 0.88 0.75 Table 62. F1 Scores across non-finetuned models for different protected categories in Tamil High-Quality Test Cases. Manuscript submitted to ACM 90Ng et al. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Ethnicity/Race/Origin 0.51 0.54 0.68 0.75 0.67 0.63 0.71 0.80 0.74 Gender/Sexuality 0.42 0.56 0.62 0.69 0.67 0.61 0.51 0.87 0.74 Religion 0.67 0.65 0.79 0.80 0.75 0.68 0.77 0.83 0.79 Table 63. F1 Scores across fine-tuned models for different protected categories in Indonesian High-Quality Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Disability 0.56 0.59 0.59 0.65 0.58 0.61 0.61 0.73 0.68 Ethnicity/Race/Origin 0.66 0.66 0.70 0.74 0.69 0.67 0.67 0.81 0.86 Gender/Sexuality 0.69 0.69 0.72 0.78 0.74 0.74 0.67 0.86 0.88 PLHIV 0.60 0.66 0.65 0.76 0.73 0.71 0.67 0.86 0.88 Religion 0.67 0.69 0.70 0.78 0.72 0.68 0.71 0.80 0.83 Table 64. F1 Scores across fine-tuned models for different protected categories in Tagalog High-Quality Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Age 0.50 0.60 0.69 0.75 0.69 0.59 0.53 0.72 0.59 Disability 0.62 0.78 0.83 0.77 0.72 0.69 0.83 0.90 0.81 Ethnicity/Race/Origin 0.65 0.65 0.75 0.73 0.70 0.65 0.74 0.79 0.74 Gender/Sexuality 0.68 0.66 0.75 0.72 0.72 0.66 0.74 0.83 0.77 Religion 0.72 0.72 0.77 0.78 0.75 0.73 0.79 0.85 0.82 Vulnerable Workers 0.63 0.60 0.77 0.75 0.67 0.57 0.72 0.85 0.76 Table 65. F1 Scores across fine-tuned models for different protected categories in Thai High-Quality Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Age 0.79 0.77 0.83 0.85 0.82 0.82 0.80 0.89 0.86 Disability 0.75 0.77 0.74 0.76 0.80 0.83 0.77 0.88 0.82 Ethnicity/Race/Origin 0.82 0.79 0.80 0.83 0.84 0.84 0.81 0.90 0.85 Gender/Sexuality 0.80 0.77 0.80 0.81 0.79 0.82 0.80 0.87 0.86 PLHIV 0.83 0.79 0.80 0.85 0.83 0.87 0.81 0.91 0.90 Religion 0.80 0.80 0.80 0.82 0.81 0.82 0.79 0.88 0.85 Table 66. F1 Scores across fine-tuned models for different protected categories in Vietnamese High-Quality Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia91 p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Age 0.46 0.39 0.64 0.63 0.52 0.49 0.69 0.67 0.58 Disability 0.50 0.47 0.65 0.68 0.64 0.57 0.63 0.75 0.61 Ethnicity/Race/Origin 0.73 0.63 0.69 0.76 0.75 0.75 0.75 0.81 0.75 Gender/Sexuality 0.63 0.59 0.75 0.78 0.75 0.65 0.70 0.83 0.74 Religion 0.64 0.69 0.77 0.80 0.78 0.73 0.74 0.80 0.75 Table 67. F1 Scores across fine-tuned models for different protected categories in Malay High-Quality Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Age 0.56 0.58 0.66 0.67 0.64 0.59 0.56 0.66 0.67 Disability 0.68 0.61 0.63 0.65 0.69 0.65 0.66 0.69 0.69 Ethnicity/Race/Origin 0.60 0.58 0.66 0.67 0.67 0.67 0.62 0.65 0.62 Gender/Sexuality 0.62 0.60 0.64 0.62 0.66 0.64 0.63 0.64 0.65 Religion 0.60 0.58 0.64 0.65 0.66 0.61 0.62 0.63 0.61 Table 68. F1 Scores across fine-tuned models for different protected categories in Mandarin High-Quality Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Age 0.51 0.52 0.78 0.84 0.53 0.47 0.59 0.87 0.66 Disability 0.66 0.58 0.74 0.73 0.68 0.71 0.76 0.77 0.63 Ethnicity/Race/Origin 0.82 0.66 0.80 0.86 0.79 0.77 0.80 0.87 0.75 Gender/Sexuality 0.76 0.65 0.78 0.83 0.76 0.73 0.79 0.88 0.79 Religion 0.84 0.74 0.84 0.88 0.83 0.79 0.83 0.91 0.84 Table 69. F1 Scores across fine-tuned models for different protected categories in Singlish High-Quality Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Age 0.62 0.61 0.70 0.77 0.61 0.59 0.66 0.76 0.76 Disability 0.63 0.60 0.63 0.78 0.56 0.60 0.63 0.81 0.83 Ethnicity/Race/Origin 0.69 0.69 0.69 0.81 0.69 0.73 0.72 0.84 0.83 Gender/Sexuality 0.73 0.71 0.64 0.78 0.67 0.69 0.71 0.82 0.83 Religion 0.76 0.75 0.72 0.81 0.72 0.73 0.69 0.86 0.86 Table 70. F1 Scores across fine-tuned models for different protected categories in Tamil High-Quality Test Cases. Manuscript submitted to ACM 92Ng et al. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Ethnicity/Race/Origin 0.70 0.69 0.56 0.60 0.77 0.67 0.60 0.75 0.76 0.76 0.83 0.60 Gender/Sexuality 0.56 0.61 0.45 0.53 0.67 0.54 0.54 0.59 0.60 0.52 0.67 0.45 Religion 0.75 0.72 0.70 0.73 0.77 0.75 0.72 0.82 0.82 0.79 0.85 0.62 Table 71. F1 Scores across non-finetuned models for different protected categories in Indonesian Silver Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Disability 0.62 0.53 0.52 0.51 0.64 0.45 0.50 0.64 0.67 0.52 0.71 0.43 Ethnicity/Race/Origin 0.69 0.66 0.66 0.63 0.71 0.61 0.62 0.73 0.75 0.66 0.76 0.54 Gender/Sexuality 0.68 0.66 0.68 0.66 0.74 0.60 0.67 0.70 0.73 0.62 0.77 0.52 PLHIV 0.67 0.66 0.68 0.74 0.72 0.66 0.81 0.75 0.73 0.70 0.71 0.63 Religion 0.69 0.68 0.61 0.58 0.69 0.69 0.64 0.68 0.68 0.64 0.74 0.54 Table 72. F1 Scores across non-finetuned models for different protected categories in Tagalog Silver Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Age 0.59 0.60 0.57 0.48 0.58 0.26 0.55 0.57 0.61 0.58 0.74 0.38 Disability 0.71 0.82 0.74 0.79 0.74 0.59 0.82 0.82 0.81 0.79 0.84 0.78 Ethnicity/Race/Origin 0.69 0.69 0.65 0.58 0.69 0.49 0.71 0.74 0.75 0.74 0.76 0.59 Gender/Sexuality 0.67 0.68 0.68 0.66 0.66 0.42 0.70 0.69 0.70 0.68 0.82 0.51 Religion 0.71 0.73 0.70 0.68 0.67 0.54 0.74 0.78 0.75 0.77 0.81 0.64 Vulnerable Workers 0.66 0.63 0.53 0.55 0.59 0.38 0.64 0.79 0.75 0.64 0.56 0.50 Table 73. F1 Scores across non-finetuned models for different protected categories in Thai Silver Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Age 0.64 0.66 0.52 0.41 0.78 0.40 0.58 0.57 0.59 0.51 0.80 0.38 Disability 0.81 0.80 0.71 0.73 0.83 0.68 0.81 0.85 0.84 0.81 0.87 0.66 Ethnicity/Race/Origin 0.83 0.81 0.74 0.72 0.82 0.71 0.83 0.85 0.85 0.84 0.87 0.74 Gender/Sexuality 0.78 0.76 0.68 0.68 0.82 0.65 0.74 0.75 0.74 0.66 0.82 0.53 PLHIV 0.82 0.84 0.81 0.86 0.82 0.78 0.83 0.87 0.83 0.77 0.84 0.83 Religion 0.83 0.84 0.80 0.78 0.83 0.78 0.85 0.85 0.86 0.86 0.89 0.74 Table 74. F1 Scores across non-finetuned models for different protected categories in Vietnamese Silver Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia93 p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Age 0.53 0.62 0.44 0.33 0.65 0.40 0.38 0.43 0.46 0.68 0.75 0.28 Disability 0.72 0.70 0.62 0.52 0.71 0.66 0.56 0.76 0.78 0.77 0.78 0.64 Ethnicity/Race/Origin 0.65 0.68 0.69 0.71 0.65 0.67 0.67 0.73 0.74 0.76 0.75 0.72 Gender/Sexuality 0.69 0.64 0.64 0.64 0.68 0.60 0.68 0.69 0.71 0.73 0.78 0.62 Religion 0.72 0.71 0.68 0.60 0.71 0.68 0.75 0.77 0.76 0.79 0.81 0.65 Table 75. F1 Scores across non-finetuned models for different protected categories in Malay Silver Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Age 0.49 0.57 0.55 0.35 0.69 0.36 0.55 0.45 0.46 0.74 0.80 0.37 Disability 0.71 0.72 0.67 0.57 0.76 0.67 0.65 0.69 0.72 0.66 0.73 0.52 Ethnicity/Race/Origin 0.73 0.72 0.71 0.64 0.74 0.71 0.72 0.71 0.71 0.76 0.80 0.58 Gender/Sexuality 0.72 0.69 0.69 0.63 0.74 0.62 0.69 0.67 0.70 0.74 0.77 0.57 Religion 0.68 0.67 0.66 0.60 0.74 0.68 0.68 0.66 0.65 0.72 0.74 0.54 Table 76. F1 Scores across non-finetuned models for different protected categories in Mandarin Silver Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Age 0.32 0.65 0.41 0.23 0.60 0.26 0.42 0.30 0.37 0.65 0.74 0.24 Disability 0.75 0.75 0.69 0.63 0.62 0.52 0.80 0.73 0.78 0.77 0.79 0.77 Ethnicity/Race/Origin 0.75 0.79 0.70 0.62 0.68 0.67 0.76 0.78 0.78 0.84 0.82 0.72 Gender/Sexuality 0.64 0.70 0.64 0.62 0.66 0.47 0.63 0.64 0.65 0.77 0.77 0.61 Religion 0.72 0.77 0.71 0.63 0.69 0.63 0.76 0.73 0.74 0.85 0.84 0.64 Table 77. F1 Scores across non-finetuned models for different protected categories in Singlish Silver Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Gemini o3 Deepseek Age 0.48 0.55 0.64 0.41 0.65 0.28 0.42 0.43 0.55 0.68 0.77 0.32 Disability 0.63 0.64 0.61 0.61 0.60 0.48 0.51 0.69 0.79 0.80 0.83 0.62 Ethnicity/Race/Origin 0.65 0.66 0.62 0.77 0.54 0.50 0.54 0.77 0.77 0.77 0.82 0.60 Gender/Sexuality 0.57 0.63 0.52 0.68 0.58 0.47 0.53 0.68 0.75 0.76 0.73 0.62 Religion 0.62 0.63 0.55 0.72 0.53 0.53 0.54 0.76 0.76 0.77 0.81 0.67 Table 78. F1 Scores across non-finetuned models for different protected categories in Tamil Silver Test Cases. Manuscript submitted to ACM 94Ng et al. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Ethnicity/Race/Origin 0.58 0.56 0.60 0.70 0.67 0.61 0.67 0.70 0.64 Gender/Sexuality 0.48 0.44 0.49 0.54 0.56 0.48 0.50 0.56 0.51 Religion 0.65 0.68 0.71 0.77 0.72 0.68 0.74 0.76 0.67 Table 79. F1 Scores across fine-tuned models for different protected categories in Indonesian Silver Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Disability 0.62 0.68 0.64 0.70 0.69 0.67 0.61 0.69 0.69 Ethnicity/Race/Origin 0.72 0.73 0.76 0.78 0.75 0.73 0.74 0.76 0.75 Gender/Sexuality 0.77 0.69 0.77 0.80 0.79 0.75 0.72 0.81 0.80 PLHIV 0.78 0.75 0.68 0.78 0.74 0.68 0.76 0.74 0.74 Religion 0.73 0.73 0.73 0.76 0.74 0.72 0.72 0.76 0.75 Table 80. F1 Scores across fine-tuned models for different protected categories in Tagalog Silver Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Age 0.41 0.55 0.58 0.60 0.54 0.38 0.52 0.53 0.53 Disability 0.56 0.74 0.86 0.81 0.70 0.65 0.78 0.78 0.79 Ethnicity/Race/Origin 0.60 0.65 0.74 0.75 0.72 0.67 0.72 0.71 0.70 Gender/Sexuality 0.63 0.64 0.72 0.73 0.72 0.60 0.65 0.68 0.65 Religion 0.65 0.70 0.76 0.77 0.74 0.71 0.75 0.74 0.71 Vulnerable Workers 0.63 0.57 0.76 0.77 0.70 0.50 0.65 0.79 0.65 Table 81. F1 Scores across fine-tuned models for different protected categories in Thai Silver Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Age 0.78 0.80 0.82 0.82 0.83 0.78 0.82 0.78 0.82 Disability 0.83 0.83 0.85 0.88 0.84 0.85 0.82 0.89 0.88 Ethnicity/Race/Origin 0.85 0.84 0.85 0.87 0.87 0.86 0.85 0.89 0.88 Gender/Sexuality 0.84 0.83 0.84 0.86 0.85 0.83 0.84 0.83 0.84 PLHIV 0.84 0.82 0.84 0.82 0.87 0.83 0.82 0.87 0.89 Table 82. F1 Scores across fine-tuned models for different protected categories in Vietnamese Silver Test Cases. Manuscript submitted to ACM SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia95 p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Age 0.42 0.37 0.64 0.59 0.48 0.48 0.62 0.45 0.44 Disability 0.68 0.53 0.74 0.75 0.67 0.69 0.77 0.70 0.64 Ethnicity/Race/Origin 0.74 0.66 0.68 0.74 0.71 0.71 0.73 0.75 0.73 Gender/Sexuality 0.70 0.63 0.78 0.76 0.73 0.68 0.75 0.75 0.72 Religion 0.66 0.70 0.77 0.79 0.74 0.73 0.77 0.78 0.72 Table 83. F1 Scores across fine-tuned models for different protected categories in Silver Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Age 0.59 0.65 0.76 0.78 0.65 0.61 0.65 0.66 0.66 Disability 0.69 0.69 0.70 0.73 0.72 0.70 0.67 0.73 0.71 Ethnicity/Race/Origin 0.65 0.63 0.72 0.69 0.70 0.66 0.66 0.67 0.61 Gender/Sexuality 0.70 0.69 0.73 0.73 0.73 0.70 0.69 0.72 0.69 Religion 0.66 0.65 0.71 0.71 0.69 0.66 0.64 0.67 0.62 Table 84. F1 Scores across fine-tuned models for different protected categories in Silver Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Age 0.45 0.40 0.72 0.73 0.64 0.46 0.57 0.75 0.75 Disability 0.63 0.57 0.77 0.74 0.70 0.67 0.76 0.77 0.71 Ethnicity/Race/Origin 0.79 0.75 0.83 0.83 0.80 0.79 0.79 0.83 0.82 Gender/Sexuality 0.77 0.73 0.81 0.81 0.76 0.74 0.75 0.81 0.81 Religion 0.80 0.82 0.81 0.84 0.79 0.76 0.81 0.83 0.81 Table 85. F1 Scores across fine-tuned models for different protected categories in Silver Test Cases. p_category Ministral Llama3b Llama8b Sealion Seallm Pangea Qwen Gemma Seagem Age 0.56 0.53 0.61 0.72 0.52 0.53 0.58 0.75 0.70 Disability 0.67 0.64 0.64 0.78 0.52 0.54 0.59 0.81 0.83 Ethnicity/Race/Origin 0.66 0.65 0.64 0.76 0.65 0.67 0.65 0.78 0.75 Gender/Sexuality 0.70 0.67 0.60 0.76 0.61 0.63 0.64 0.80 0.79 Religion 0.71 0.68 0.65 0.74 0.68 0.68 0.65 0.78 0.76 Table 86. F1 Scores across fine-tuned models for different protected categories in Silver Test Cases. Manuscript submitted to ACM