Paper deep dive
SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing
Yifan Lyu, Dianqing Lin, Xinran Li, Jiaqi Qiao, Xiujuan Xu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/28/2026, 2:57:33 AM
Summary
The paper introduces SPAR-Hate, an auditor-guided multi-perspective role-reasoning framework for bilingual hate speech parsing. It addresses challenges in structured parsing, such as shared-target multi-tuple binding and culturally coded slurs, by decomposing documents into local focus units and generating evidence-grounded candidates from Victim, Moderator, and Cultural Bystander perspectives. These candidates are clustered and arbitrated under schema constraints before reassembly. Experiments on STATE-ToxiCN (Chinese) and TBO (English) demonstrate improved performance in joint target-argument-label metrics compared to baselines.
Entities (10)
Relation Signals (9)
SPAR-Hate ā evaluatedon ā STATE-ToxiCN
confidence 95% Ā· Experiments on STATE-ToxiCN and a controlled TBO split show gains
SPAR-Hate ā evaluatedon ā TBO
confidence 95% Ā· Experiments on STATE-ToxiCN and a controlled TBO split show gains
SPAR-Hate ā usesperspective ā Cultural Bystander
confidence 95% Ā· SPAR-Hate ... elicits evidence-grounded candidates from Victim, Moderator, and Cultural Bystander perspectives
SPAR-Hate ā usesperspective ā Victim
confidence 95% Ā· SPAR-Hate ... elicits evidence-grounded candidates from Victim, Moderator, and Cultural Bystander perspectives
SPAR-Hate ā usesperspective ā Moderator
confidence 95% Ā· SPAR-Hate ... elicits evidence-grounded candidates from Victim, Moderator, and Cultural Bystander perspectives
Cultural Bystander ā handles ā Culturally Coded Slurs
confidence 92% Ā· Cultural Bystander emphasises local cultural context, community slang, pragmatic history, and coded expression.
SPAR-Hate ā backedby ā Qwen3-14B
confidence 90% Ā· The main local experiments use Qwen3-14B
SPAR-Hate ā backedby ā DeepSeek V4 Pro
confidence 90% Ā· SPAR-DeepSeek-v4-Pro56.80 65.96
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Hate speech research has moved from coarse-grained classification towards structured parsing, where systems jointly identify targets, supporting arguments, and target-level labels. Documents with multiple targets, conflicting local readings, or culturally coded language make these bindings difficult to recover. SPAR-Hate is an auditor-guided multi-perspective role-reasoning framework for bilingual hate speech parsing. It decomposes each document into local focus units, elicits evidence-grounded candidates from Victim, Moderator, and Cultural Bystander perspectives, resolves candidate conflicts under grounding and schema constraints, and reassembles sample-level predictions. Experiments on STATE-ToxiCN and a controlled TBO split show gains across local and API backbones, concentrated on strict joint target-argument-label metrics. Full-test integrated-prompt controls, component ablations, and bounded-arbitration diagnostics identify the contribution of separated perspective generation and arbitration. Structured teacher traces also support training a smaller student model.
Tags
Links
- Source: https://arxiv.org/abs/2608.22018v3
- Canonical: https://arxiv.org/abs/2608.22018v3
Trouble viewing inline? Open PDF directly ā
Full Text
68,928 characters extracted from source content.
Expand or collapse full text
SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing Yifan Lyu 1 Dianqing Lin 2 Xinran Li 1 Jiaqi Qiao 1 Xiujuan Xu 1,* 1 Dalian University of Technology 2 Inner Mongolia University stevelyu811@gmail.com * Correspondence:xjxu@dlut.edu.cn Abstract ĆWarning: This paper contains content that may be offensive or harmful. Hate speech research has moved from coarse- grained classification towards structured pars- ing, where systems jointly identify targets, supporting arguments, and target-level la- bels. Documents with multiple targets, con- flicting local readings, or culturally coded language make these bindings difficult to recover. SPAR-Hate is an auditor-guided multi-perspective role-reasoning framework for bilingual hate speech parsing. It de- composes each document into local focus units, elicits evidence-grounded candidates from Victim, Moderator, and Cultural By- stander perspectives, resolves candidate con- flicts under grounding and schema constraints, and reassembles sample-level predictions. Ex- periments on STATE-ToxiCN and a controlled TBO split show gains across local and API backbones, concentrated on strict joint targetā argumentālabel metrics. Full-test integrated- prompt controls, component ablations, and bounded-arbitration diagnostics identify the contribution of separated perspective genera- tion and arbitration. Structured teacher traces also support training a smaller student model. 1 Introduction Hate speech detection identifies hateful expres- sions whose interpretation depends on context, so- cial judgement, and platform standards (Schmidt and Wiegand,2017;Fortuna and Nunes,2018). Fine-grained work now includes rationale anno- tation, targetāargument parsing, and tuple extrac- tion ( Pavlopoulos et al.,2021;Mathew et al.,2021; Zampieri et al.,2023;Bai et al.,2025). Structured hate parsing predicts one or more target-level tu- EN Challenge 1: Shared-Target Multi-Tuple Binding INPUT All you senatorssuck as human beings you are not protecting the citizens but playing politics. Gold tuples (senators, suck as human beings, 1) (senators, not protecting the citizens, 1) Typical parsing error (senators, playing politics, 0) Under a shared target, one gold tuple is omitted, and the retained prediction misbinds the target to a nearby non-gold span and assigns an incorrect label. ZH Challenge 2: Coded and Culturally Localised Slurs INPUT é¾čåäøęø ę„čŖå·±ę仄å¤ēå ¶ä»å°åŗļ¼ äøå¾å½äøŗäøåäŗŗć Gold tuple (é¾č, åäøęø ę„čŖå·±ę仄å¤ēå ¶ä»å°åŗ, Region, 1) Typical parsing error (é¾č, åäøęø ę„čŖå·±ę仄å¤ēå ¶ä»å°åŗ, Region, 0) Despite correct span recovery, failure to recognise āé¾čā as a homophonic regional slur yields a neutral label. Figure 1: Typical LLM errors in bilingual hate parsing, including omitted tuples, imprecise targetāargument binding, and failed interpretation of culturally localised homophonic slurs. ples, each linking an attacked target, its supporting argument, and a target-level label. Figure 1shows omitted tuples, faulty targetā argument binding, and a missed homophonic slur. These failures require consistent tuple prediction and culturally grounded interpretation. Large lan- guage models remain skewed towards Western cul- tural representations ( Naous et al.,2024), weak- ening performance on implicit, coded, and lo- cally grounded hate (ElSherief et al.,2021;Nozza, arXiv:2608.22018v3 [cs.AI] 27 Aug 2026 2021;Ocampo et al.,2023;Lin et al.,2026). Cultural knowledge can change target identifica- tion, evidence selection, and harm attribution. So- cial position also shapes interpretations of hostile speech (Spears,2021), so the same utterance can support different grounded readings. SPAR-Hate comprises four stages: Segment, Perspective-Guided Role Generation, Arbitrate, and Reassemble. Phase 1 produces local focus units. Victim, Moderator, and Cultural Bystander generators produce evidence-grounded candidates for each unit. Phase 3 clusters and selects can- didates under grounding and schema constraints, and Phase 4 restores sample-level predictions. The Chinese Cultural Bystander receives weak lexical context for coded language and local slang ( Bai et al. ,2025;Xiao et al.,2024;Lu et al.,2023). Across local and API settings, SPAR-Hate im- proves the stricter joint structural metrics on Chi- nese STATE-ToxiCN and the controlled English TBO split. Component ablations test each phase, integrated-prompt controls test separated perspec- tive generation and arbitration, and distillation transfers the structured traces to a smaller student model. Contributions.The task formulation aligns Chi- nese and English structured parsing at field level while retaining their original annotation contracts. The framework combines local focus-unit segmen- tation, three perspective-conditioned generators, constrained arbitration, and sample-level reassem- bly. Main experiments, full-test integrated-prompt controls, ablations, and diagnostics test strict struc- tural recovery and structured teacher-trace trans- fer. 2 Related Work 2.1 From Hate Speech Classification to Structured Tuple Parsing Hate speech research began largely as text clas- sification, but class labels alone do not support fine-grained semantic understanding. Early work focused on definitions, features, and classifica- tion models ( Schmidt and Wiegand,2017;For- tuna and Nunes,2018;Waseem and Hovy,2016; Davidson et al.,2017). Later multilingual shared tasks and richer annotations showed that multi- lingual hate analysis requires more than single- label prediction ( Basile et al.,2019;Ousidhoum et al.,2019). Toxic Spans, HateXplain, TBO, and STATE-ToxiCN then moved the field toward fine- grained localisation, rationale annotation, target- argument parsing, and Chinese quadruple parsing (Pavlopoulos et al.,2021;Mathew et al.,2021; Zampieri et al.,2023;Bai et al.,2025). SPAR-Hate focuses on local targetāargumentā label bindings under separate Chinese and English annotation contracts and evaluates their sample- level reconstruction. 2.2 Implicit Hate, Cultural Context, and Cross-Lingual Fragility Implicit, subtle, and context-dependent hate re- mains difficult to detect. Coded and indirect ex- pressions weaken model performance ( ElSherief et al.,2021;Ocampo et al.,2023;Hartvigsen et al., 2022), while HateCheck identifies related func- tional failures (Rƶttger et al.,2021). Cross-lingual zero-shot models can also misread language- specific non-hateful taboo expressions as hate sig- nals (Nozza,2021). Explainable hate-speech detection uses ratio- nales, social bias frames, and stepwise expla- nations ( Mathew et al.,2021;Sap et al.,2020; Yang et al.,2023). Structured parsing adds the requirement that target, argument, and label re- main jointly aligned. SPAR-Hate supplies cultural knowledge as weak context and retains only text- grounded candidates. 2.3 Role-Conditioned Reasoning and Constrained Aggregation RoleLLM and multi-perspective role-playing show that role conditioning elicits distinct knowl- edge and reasoning biases (Wang et al.,2024). Proposerāaggregator methods coordinate several model outputs (Du et al.,2023). AutoGen and MetaGPT organise role allocation and structured workflow handoffs, with MetaGPT encoding these handoffs through standard operating procedures ( Wu et al.,2023;Hong et al.,2024). SPAR-Hate constrains role outputs to compa- rable targetāargumentālabel candidates before ag- gregation. Its distillation traces retain focus-unit decomposition, role hypotheses, conflict diagno- sis, and final arbitration, linking the method to chain-of-thought, self-consistency, and reasoning distillation ( Wei et al.,2023;Wang et al.,2023; Shridhar et al.,2023;Hsieh et al.,2023). 3 Methods: The SPAR-Hate Framework SPAR-Hate combines local focus-unit decomposi- tion, multi-perspective candidate generation, dy- namic arbitration, and sample-level reassembly in an auditor-guided framework for bilingual hate parsing. The pipeline breaks document-level pars- ing into explicit intermediate stages, which helps long texts, multi-target cases, implicit attacks, and culturally coded slang. 3.1 Task Formulation and Output Contracts Given an input documentD, the system must pre- dict a set of target-level structured tuples E=e 1 ,e 2 ,...,e n .(1) The Chinese and English benchmarks follow different original annotation contracts. STATE- ToxiCN uses quadruples e raw zh = (target,argument,group,hateful), (2) wheregroupdenotes the attacked group category. TBO uses triples e raw en = (target,argument,harmful).(3) These schemas correspond respectively to the TargetāArgumentāHatefulāGroup annotation con- tract in STATE-ToxiCN and the targetāargumentā harmfulness contract in TBO ( Bai et al.,2025; Zampieri et al.,2023). To construct the bilingual main track at field level, the Chinese main track removesgroupand uses a three-field output e main zh = (target,argument,label).(4) For unified notation, each tuple on the aligned bilingual main tracks is written ase i = (t i ,a i ,ā i ). Heret i denotes the attacked target,a i the attack argument, andā i ā0,1the target-level harmful- ness label. It corresponds toharmfulin English andhatefulin Chinese. The Chinese four-field extension retainsgroupas an additional field. The Chinese three-field task supports the bilin- gual main comparison. A four-field extension re- tainsgroupand tests adaptation to the original language-specific schema. Separate prompt tem- plates and output constraints preserve both con- tracts. 3.2 Framework Overview SPAR-Hate has four stages:Segment,Perspective- Guided Role Generation,Arbitrate, andReassem- ble, corresponding to Phase 1ā4. The system splits a document into local focus units, generates structured candidates from three perspectives, arbi- trates them with a dynamic-beacon procedure, and reassembles benchmark-aligned sample-level out- puts. Figure2presents the full pipeline. 3.3 Phase 1: Local Focus-Unit Segmentation Long texts often contain multiple targets, local stances, and interfering attack fragments. Di- rect extraction from the full document can there- fore produce targetāargument mismatches, overly wide arguments, and merged events. Divide-and- conquer frameworks for document-level sentiment parsing suggest the same advantage for structured extraction ( Wang et al.,2026). Phase 1 therefore uses an LLM-based segmenter to map document Dto a set of local decision units G=g 1 ,g 2 ,...,g k .(5) Each local unitg i stores sample-level and lo- cal indices (group_id,local_group_id), local text (group_text), a focus-marked local con- text (focus_text), and a soft target anchor (canonical_target). A focus unit is defined around a target and local intent while retain- ing enough context for argument grounding. Its boundary need not coincide with a syntactic boundary. 3.4 Phase 2: Perspective-Guided Role Generation Perspective shapes hate judgements (Waseem and Hovy,2016;Sap et al.,2020). The task-motivated role set covers affected-group harm, platform- governance boundaries, and community interpreta- tion of implicit or culturally coded hate. For each local unitg i , the framework instantiates three role- conditioned generators R=r victim ,r moderator ,r bystander .(6) Victim emphasises felt harm, exclusion, and mi- croaggressions. Moderator emphasises platform- governance boundaries and explicit rule violations. Cultural Bystander emphasises local cultural con- text, community slang, pragmatic history, and coded expression. Each role generates one or more ļ BILINGUAL INPUT [ZH] zh_2010 第äøå°č±”ļ¼ ę²³åļ¼ē©·ļ¼äŗŗå¤ļ¼ äøåäøēļ¼åŗēē å Øēåę°ē¬¬äø ę°ēļ¼ę²»å®äøå„½ļ¼ Translation: First impressions, Henan: poor, populous; the three northeastern provinces: the world's lowest birth rate; Xinjiang: public safety may not be very good. [EN]parallel Politicians feed on bigotry. same parsing schema ā SEGMENTER: PHASE 1 text First impressions, Henan: poor, populous; the three northeas... sample_id: zh_2010 g1 ... group_id: zh_2010_g1 g3 ... group_id: zh_2010_g3 g2 <focus>the three northeastern provinces: the world's lowest birth rate</focus> group_id: zh_2010_g2 ļ„ ROLE GENERATOR: PHASE 2 (g2) ļ¤ Victim HATE Ā· 1 Target: the three northeastern provinces Argument: the world's lowest birth rate Rationale: Linking the three northeastern provinces to having the world's lowest birth rate conveys a clear negative stereotype and demeans the group. ļ” Moderator NON-HATE Ā· 0 Target: the three northeastern provinces Argument: the world's lowest birth rate Rationale: The focused clause cites data as a demographic description of a specific region, contains no insulting wording, and can be read as an objective statement. ļµ Bystander HATE Ā· 1 Target: the three northeastern provinces Argument: the world's lowest birth rate Rationale: Using birth-rate data in this way reinforces a regional stereotype and constitutes structural bias against the northeastern region. ļ” AUDITOR ENGINE: PHASE 3 (g2) Quality & Signals grounded target aligned valid structure multi-role support lexicon bonus quote caution victim microaggr. moderator strong Cluster Formation c1: HATE Victim + Bystander ļ¤ ļµ Triage & Routing ā consensus conflict send to Auditor LLM ā arbitrate defective āŗ Role Regenerate Arbitration & Clause Verdicts ļ§ Auditor LLM selects c1 g2 ā HATE Auditor rationale: In the "first impressions" framing, "the world's lowest birth rate" is used as a negative generalization about the three northeastern provinces rather than a neutral statistical description. Therefore, c1 is selected. ļ§© REASSEMBLE: PHASE 4 g1 (Henan, poor, populous, Region, 1) g2 (the three northeastern provinces, the world's lowest birth rate, Region, 1) g3 (Xinjiang, public safety may not be very good, Region, 1) ā MERGE ā Multi-Target Restored ("Henan", "poor, populous", "Region", hate ("the three northeastern provinces", "the world's lowest birth rate", "Region", hate ("Xinjiang", "public safety may not be very good", "Region", hate c2: NON-HATE Moderator ļ” Figure 2: Overview of SPAR-Hate. Bilingual input is segmented into local focus units, processed by three role- conditioned generators, arbitrated under evidence constraints, and reassembled into sample-level outputs. Distilla- tion uses the earlier phases to construct teacher traces for student training. structured candidates over the samefocus_text, yielding the role-specific candidate set H r (g i ) =h (1) r ,h (2) r ,....(7) SPAR-Hate imposes a strict JSON schema and requiresargumentto be a supporting substring in- sidefocus_text. Every judgement includes local textual evidence. The Chinese Cultural Bystander receives entries from the auxiliary STATE-ToxiCN lexicon. Its 829term/category/definitionrecords contain no sample identifiers, gold tuples, or instance la- bels. Retrieved entries provide weak context for the currentfocus_textandcanonical_target. Candidates remain subject to focus-span ground- ing, and lexicon-only singleton candidates are dis- carded. Lexicon hits occur in 10.1% ofZH-main groups and 9.6% ofZH-quadruplegroups. Chi- nese toxicity research motivates this support for slang, homophones, and emoji cloaking ( Lu et al., 2023;Bai et al.,2025;Xiao et al.,2024). Ap- pendixBgives the prompt templates. 3.5 Phase 3: Local Candidate Clustering and Auditor Arbitration Phase 3 filters, clusters, and ranks the role candi- dates for each local unit. Given H i = āŖ rāR H r (g i ),(8) the procedure applies quality scoring, cluster for- mation, triage, and deterministic selection. Candidate Quality ScoringThe system first as- signs each candidatecāH i a quality scoreq(c): q(c) = ā j α j Ļ j (c).(9) HereĻ j (c)includes grounding, target explicitness, consistency withcanonical_target, structural completeness, label validity, and, in the Chinese four-field setting,groupālabel coherence. Chi- nese bystander candidates can receive a small lex- icon bonus. Candidates below quality0.60are fil- tered before clustering. An ungrounded or missing argument fails validation. Cluster FormationCandidates are then clustered by similarity over target, argument, harmful/hateful, and, in the Chinese four-field extension,group. Similarity between a candidate and a cluster is written as sim(c,C) =β t s t (c,C) +β a s a (c,C) +β y s y (c,C) +β g s g (c,C), (10) wheres t ,s a ,s y ,s g denote similarity intarget, argument,label, andgroup. On the three-field main tracks,β g = 0. The reported field weights are 0 . 45 / 0 . 45 / 0 . 05 / 0 . 05 , respectively. Each clus- ter is then compressed into a canonical tuple and assigned a cluster score based on member quality, the number of supporting roles, and any lexicon- aware bonus: Score(C) = ā cāC q(c)+Ī» 1 |supp(C)|+Ī» 2 I lex (C), (11) wheresupp(C)is the set of roles supporting that cluster. The mainline support-role and lexicon- cluster bonuses are0.12and0.08; the candidate- level lexicon bonus is0.05. Triage and Dynamic Soft BeaconsAfter clus- tering, the system routes each local unit into one of three lanes according to role agreement, the margin of the top cluster, and signals of structural failure or missing roles. A clearly dominant top cluster with no obvious defect entersconsensus. Missing roles or a lack of valid grounded candi- dates entersdefective. The remaining cases en- terconflict. The system then applies grounding- aware, lexicon-aware, and quote-aware soft bea- cons to adjust cluster ranking. If some roles are marked defective and regeneration is enabled, one bounded regeneration round is allowed. Highly conflicting cases can trigger a second auditor re- finement. Deterministic Fallback and Local Refinement Final ranking uses a deterministic cluster-level fallback. The selected primary cluster undergoes local refinement oftargetandargumentbound- aries. The reported configuration retains at most two clusters, requires scoreā„1.05and at least two supporting roles, and allows one regeneration round for a defective role. Across retained runs, 55.9ā61.7% of focus units enter the conflict lane and 65.1ā70.7% invoke auditor refinement. Ap- pendix Dreports the complete scoring, routing, and boundedness diagnostics. 3.6 Phase 4: Sample-Level Multi-Target Reassembly In this task setting, local focus-unit decisions are made internally while the benchmark expects sample-level predictions. Phase 4 therefore re- stores earlier decisions to the output format re- quired by evaluation. The system first aggregates all local tuples bysample_id, producing Ģ E(x) = ā g i āG(x) Y(g i ),(12) whereY(g i )is the set of final tuples output by Phase 3 for local unitg i . The main configura- tion uses a concatenate-first strategy, preserving local provenance and not forcing exact deduplica- tion. The default exported sample-level prediction is therefore Ė E(x) = Ģ E(x),(13) whereas under optional exact deduplication Ė E(x) = Dedup ( Ģ E(x) ) .(14) 4 Experiments 4.1 Datasets and Evaluation Protocol DatasetsExperiments are conducted on Chinese STATE-ToxiCN and English TBO (Bai et al., 2025;Zampieri et al.,2023). STATE-ToxiCN keeps its official train/test split. The public re- lease of TBO provides only a test split, so the public test set is deterministically shuffled with seed=20260415and re-divided into 3200/800 train/test subsets. Beyond this repartition, process- ing is limited to field cleaning, schema normalisa- tion, and split freezing. No extra relabelling is in- troduced, and multi-target documents are not split at this stage. The paper reports three evaluation tracks,ZH-main,EN-main, andZH-quadruple, all defined exactly as in Section 3.1. Dataset compo- sition appears in Table 1. TBO numbers are con- trolled within-paper comparisons on this determin- istic split only; they are not directly comparable to studies evaluated on the original public test-only release. Evaluation metricsThe original evaluation pro- tocols of both benchmarks are kept. No mixed cross-lingual total score is constructed. For Chi- nese, following STATE-ToxiCN, both Hard and Soft Macro-F1 are reported: Hard requires ex- act agreement with gold spans and field combina- tions, while Soft credits overlapping spans under the same target-centred structure ( Bai et al.,2025). DatasetPublic SplitFinal Split Tracks STATE-ToxiCN official train/test 6424 / 1605 ZH-main, ZH-quadruple TBOpublic test only 3200 / 800 EN-main Table 1: Datasets used in the experiments. TBO is repartitioned from its public test set with seed=20260415. API zero-shot model EN ZH DeepSeek-v4-Pro23.18 22.80 GPT-5.421.01 28.13 Table 2: Full-test average scores for the added API zero-shot baselines. (a)ZH-main ModelTargetArgumentT-A PairT-A-H Tri.Avg. Hard Soft Hard Soft Hard Soft Hard Soft Local LLMs 0-shot-Qwen3-14B45.49 54.9716.24 46.99 10.62 32.73 7.65 23.93 29.83 SPAR-Qwen3-14B (mainline)48.73 55.75 21.0356.8512.3634.148.7124.9732.82 SPAR-Qwen3.5-35B45.8453.12 20.0957.9511.2335.87 9.26 27.2432.58 API Models few-shot-DeepSeek-v4-Pro52.69 63.08 17.86 61.03 13.62 39.53 10.75 31.09 36.21 few-shot-GPT-5.454.1765.8318.8762.1013.82 43.9010.60 31.47 37.60 SynChain-DeepSeek-v4-Pro38.92 48.68 11.54 49.57 5.09 31.89 3.75 23.75 26.65 SynChain-GPT-5.434.89 43.15 11.52 45.63 5.38 30.28 4.04 21.98 24.61 Dance-DeepSeek-v4-Pro53.19 62.77 19.34 57.40 15.75 41.52 12.55 32.47 36.87 Dance-GPT-5.451.99 60.73 23.8957.7219.3244.5613.6333.9938.23 SPAR-DeepSeek-v4-Pro56.80 65.96 24.0362.0617.27 44.13 13.01 32.8839.52 SPAR-GPT-5.452.62 61.94 23.12 56.62 17.4245.4813.3234.3838.11 (b)EN-main ModelTarget Argument Targeted NTA TA Harm Avg. Local LLMs 0-shot-Qwen3-14B32.4146.4825.0614.91 11.06 47.61 29.59 SPAR-Qwen3-14B (mainline) 43.4547.9634.6820.96 18.6859.3137.51 SPAR-Qwen3.5-35B46.4950.2235.4320.2816.7260.11 38.21 API Models few-shot-DeepSeek-v4-Pro36.2642.4022.123.09 1.73 53.98 26.60 few-shot-GPT-5.438.9444.0213.504.00 2.61 48.16 25.21 SynChain-DeepSeek-v4-Pro38.7041.2218.262.85 2.5858.0026.94 SynChain-GPT-5.436.6037.5215.564.76 3.52 51.92 24.98 Dance-DeepSeek-v4-Pro45.6736.0827.432.92 2.61 56.44 28.53 Dance-GPT-5.446.15 42.3924.928.88 7.75 52.62 30.45 SPAR-DeepSeek-v4-Pro51.3150.3630.0219.8113.3356.1936.84 SPAR-GPT-5.442.1547.0029.2418.8415.8156.7834.97 Table 3: Main results onZH-mainandEN-main. Within each split, the best result in each column is boldfaced and the second-best is underlined. On the Chinese tracks, Target and Argument evaluate target and argument field parsing. T- A Pair requires the target to be correctly bound to its argument. T-A-H Tri and Quad require the local structure to remain correct after adding the hateful label and, in the four-field task, the groupfield. For English, the paper follows TBOās tuple-level evaluation: Target and Argument mea- sure field-level parsing; Targeted requires joint parsing of target and harmfulness; NTA and TA require exact match on(target,argument) and(target,argument,harmful)respectively; Harm evaluates harmfulness on the predicted tar- get tuples ( Zampieri et al.,2023). Chinese met- rics therefore emphasise hard/soft field parsing, whereas English metrics emphasise exact tuple consistency ( Bai et al.,2025;Zampieri et al., 2023). 4.2 Experimental Setup and Baselines Experimental setupThe main local experi- ments use Qwen3-14B under a dual-24GB-class GPU budget (Yang et al.,2025). Phase 1ā3 share the same backbone, with lexicon injection used only for the Chinese Cultural Bystander. Long runs support checkpoint-resume and OOM split-retry. Phase 4 is deterministic aggregation. Sample-level outputs follow the concatenate-first strategy of Section 3.6. AppendixAgives the run- time configuration and artifact map; AppendixD gives the arbitration settings. BaselinesThe baselines cover direct extraction, task-adapted SynChain and Dance, and SPAR with local and API backbones (Fan et al.,2025; Wang et al.,2026;DeepSeek-AI,2026;OpenAI, 2026). Table2reports the added full-test API zero-shot averages. Prompt templates and baseline adaptations appear in AppendicesBandC. 4.3 Main Results Table3reports the complete bilingual main-track comparison. Main bilingual triplet resultsUnder a fixed lo- cal 14B backbone, SPAR raises the average score onZH-mainfrom29.83to32.82and onEN-main from29.59to37.51. The larger gains occur on stricter joint structural metrics. Qwen3.5-35B records38.21onEN-main and32.58onZH-main. Among API systems, SPAR-DeepSeek-v4-Prorecords the highest ZH-mainaverage (39.52), whileSPAR-GPT-5.4 records EN NTA/TA scores of18.84/15.81. Chinese quadruple extensionThe Chinese four-field extension (ZH-quadruple) retains the groupfield as a test of adaptation to language- specific schema. Full results appear in Appendix Table25. Cross-model observationsThe largest differ- ences occur on Targeted, NTA, and TA for English and on T-A Pair, T-A-H Tri, and Quad for Chinese. Appendix Fprovides additional outputs and cases. 4.4 Framework Analysis Integrated-prompt controlsFour full-test integrated-prompt controls fix the Phase 1 focus units and place all three perspective descriptions in one call. The Chinese controls retain the same lexicon access. Table 4reports lower strict tuple-binding scores after removing separated role outputs, candidate clustering, and auditor arbitration. GPT-5.4 Target scores are46.01for integrated prompting and42.15for SPAR on EN, with corresponding ZH scores of54.40and52.62. Component ablationsTable5summarises the most diagnostic hard and exact metrics; complete results are deferred to AppendixH. Phase-wise ablationsRemoving the segmenter raises Chinese Target Hard F1 from48.73to52.34 and lowers T-A-H Tri Hard F1 from8.71to7.02. On English, Target F1 rises from43.45to46.95, while TA and Harm fall. Removing Phase 3 lowers the ZH and EN averages to31.71and36.72. Role-wise ablationsRemoving the bystander changes the Chinese average from32.82to32.85 and lowers T-A Pair Hard and T-A-H Tri Hard. Removing the moderator lowers English TA from 18.68to17.51, while NTA changes from20.96to 21.34. BackboneTrackIntegratedSPAR DeepSeek-v4-Pro EN NTA / TA 4.87 / 3.89 19.81 / 13.33 GPT-5.4EN NTA / TA 11.41 / 8.18 18.84 / 15.81 DeepSeek-v4-Pro ZH TA / TAH-H 6.31 / 4.40 17.27 / 13.01 GPT-5.4ZH TA / TAH-H 14.08 / 10.56 17.42 / 13.32 Table 4: Full-test fixed-Phase-1 integrated-prompt con- trols. The integrated condition retains the three lens descriptions (and Chinese lexicon access) but removes separated candidates, clustering, and auditor arbitra- tion. (a)ZH-main SettingTgt-H Arg-H TA-H TAH-H Avg. Qwen3-14B (mainline) 48.73 21.03 12.36 8.71 32.82 w/o P1 segmenter52.34 15.20 9.26 7.02 32.25 w/o P2 victim50.40 15.95 10.53 7.63 31.16 w/o P2 moderator49.90 16.51 10.86 7.75 31.77 w/o P2 bystander50.18 16.10 10.50 8.20 32.85 w/o P3 arbitration46.54 17.03 10.67 7.68 31.71 (b)EN-main SettingTgt NTA TA Harm Avg. Qwen3-14B (mainline) 43.45 20.96 18.68 59.31 37.51 w/o P1 segmenter46.95 20.65 18.04 57.21 37.34 w/o P2 victim44.70 21.18 17.20 57.21 37.02 w/o P2 moderator44.29 21.34 17.51 59.11 37.43 w/o P3 arbitration 43.66 20.42 17.81 57.61 36.72 Table 5: Ablation summary on the bilingual main tracks. The main text keeps only the most diagnostic hard/exact metrics; complete bilingual ablation results appear in Appendix H. Diagnostic validationSupplementary exact-set and blinded semantic diagnostics appear in Ap- pendix E. They are reported as secondary checks, not primary ranking metrics. The attacked-group breakdown in the same appendix locates a remain- ing multi-group binding error. 4.5 Distillation Results Table6compares the distilled 4B student with the zero-shot and SPAR-Qwen3-14B systems on all three evaluation tracks. The distilled student improves over SPAR- Qwen3-14B on both Chinese tracks and remains close to its teacher onEN-main, despite using a substantially smaller backbone. Appendix Gre- ports teacher-trace construction, filtering, and stu- dent training. (a)ZH-main ModelTAH-H TAH-S Avg. 0-shot-Qwen3-14B7.6523.93 29.83 SPAR-Qwen3-14B (mainline) 8.7124.97 32.82 Distillation-Qwen3-4B8.3326.04 33.67 (b)EN-main ModelTA Harm Avg. 0-shot-Qwen3-14B11.06 47.61 29.59 SPAR-Qwen3-14B (mainline) 18.68 59.31 37.51 Distillation-Qwen3-4B18.50 59.05 37.33 (c)ZH-quadruple ModelQuad-H Quad-S Avg. 0-shot-Qwen3-14B6.8719.42 25.70 SPAR-Qwen3-14B (mainline)6.1822.37 28.95 Distillation-Qwen3-4B7.0425.05 30.38 Table 6: Distillation summary across the three evalua- tion tracks. 4.6 Qualitative Analysis Case studyTable7traces three Phase 1 local focus units from a Chinese multi-target example. Phase 2 produces conflicting labels forg 2 andg 3 , and Phase 3 selects the gold label for both units. Forg 1 , SPAR retains the full argument āpoor; pop- ulousā, which SynChain truncates. ItemContent TextFirst impression: Henan is poor and populous; Northeast China has one of the worldās lowest birth rates; public se- curity in Xinjiang may not be very good. Goldg 1 : (Henan, āpoor; populousā, 1) g 2 : (Northeast China, one of the worldās lowest birth rates, 1) g 3 : (Xinjiang, public security may not be very good, 1) Phase 1g 1 : Henan: poor; populous g 2 : Northeast China: one of the worldās lowest birth rates g 3 : Xinjiang: public security may not be very good Phase 2g 1 : only hateful candidates remain. g 2 : (Northeast China, one of the worldās lowest birth rates, 1/0). g 3 : (Xinjiang, public security may not be very good, 1/0). Phase 3g 1 : (Henan, āpoor; populousā, 1) g 2 : (Northeast China, one of the worldās lowest birth rates, 1) g 3 : (Xinjiang, public security may not be very good, 1) Baselinesg 1 : Dance is correct. SynChain outputs (Henan, poor, 0), missing āpopulousā. g 2 : Dance and SynChain both assign label 0. g 3 : Dance and SynChain both assign label 0. Table 7: Phase-wise processing of three local focus units.g 1 ,g 2 , andg 3 denote Phase 1 units. One unit may yield several sample-level tuples after Phase 4. The Chinese input is shown in English translation. Error analysisThe main residual error is argu- ment boundary over-expansion. In Table8, SPAR- Hate predicts the correct target and harmfulness but retains additional contextual material, so the prediction receives no tuple credit under exact match. Appendix Fprovides further cases. ItemContent Textsome of yāall arenāt being held accountable for the dumb shit that yāall be doing....and it shows Gold(yāall, dumb shit, 1) SPAR- Qwen3-14B (yāall, dumb shit that yāall be doing, 1); (null, it shows, 0) Table 8: A remaining SPAR-Hate error: overly wide argument boundaries. 5 Conclusion SPAR-Hate combines local focus-unit segmenta- tion, role-conditioned generation, constrained au- ditor arbitration, and sample-level reassembly. Its largest gains occur on stricter joint structural metrics under the controlled benchmark settings. Integrated-prompt controls record lower tuple- binding scores after collapsing perspective gener- ation and arbitration into one call. The Chinese four-field extension and distillation results cover language-specific schema and structured teacher traces. Limitations Evaluation covers Chinese and English and omits a full comparison with task-specific fine-tuned sys- tems. The task-motivated role set has not been ex- haustively searched, and the lexicon has no com- plete removal ablation. Weak lexical context can still over-flag dialectal, culturally loaded, or re- claimed terms. The threshold replay, blind audit, and attacked-group breakdown have limited scope and cannot support broad claims about statistical optimality, fairness, or cultural validity. TBO re- sults use the deterministic internal split and do not support direct comparison with results on the orig- inal public test-only release. Ethical Considerations This study uses public Chinese and English bench- mark datasets to examine structured hate speech parsing. Their contents include abusive, discrimi- natory, and potentially traumatic expressions. The paper retains only representative examples needed for the academic argument and recommends a minimal-exposure principle in data processing, vi- sualisation, and manual analysis to reduce sec- ondary harm to researchers, annotators, and read- ers. Such systems also carry misuse risks, includ- ing over-censorship, automated punishment, and large-scale opinion monitoring. Earlier work shows that false positives can arise from am- biguous boundaries between offensive language and hate speech, spurious correlations around minority-group mentions, and model fragility on functional phenomena (Davidson et al.,2017; Hartvigsen et al.,2022;Rƶttger et al.,2021). Structured outputs and traceable intermediate pro- cesses increase transparency, but they may also in- crease the operational reach of deployed moder- ation systems. Real-world deployment therefore requires human review, appeals, threshold calibra- tion, error auditing, and continuing bias evaluation across languages, groups, and cultural settings. References Zewen Bai, Liang Yang, Shengdi Yin, Junyu Lu, Jingjie Zeng, Haohao Zhu, Yuanyuan Sun, and Hongfei Lin. 2025.STATE ToxiCN: A benchmark for span-level target-aware toxicity extraction in Chi- nese hate speech detection. InFindings of the As- sociation for Computational Linguistics: ACL 2025, pages 10206ā10219, Vienna, Austria. Association for Computational Linguistics. Valerio Basile, Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco Manuel Rangel Pardo, Paolo Rosso, and Manuela San- guinetti. 2019. SemEval-2019 task 5: Multilin- gual detection of hate speech against immigrants and women in Twitter. InProceedings of the 13th Inter- national Workshop on Semantic Evaluation, pages 54ā63, Minneapolis, Minnesota, USA. Association for Computational Linguistics. Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017.Automated Hate Speech Detection and the Problem of Offensive Language. Proceedings of the International AAAI Conference on Web and Social Media, 11(1):512ā515. DeepSeek-AI. 2026.Deepseek-v4: Towards highly ef- ficient million-token context intelligence. Technical report. Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023.Qlora: Efficient finetuning of quantized llms .Preprint, arXiv:2305.14314. Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. 2023.Improving factuality and reasoning in language models through multiagent debate .Preprint, arXiv:2305.14325. Mai ElSherief, Caleb Ziems, David Muchlinski, Vaish- navi Anupindi, Jordyn Seybolt, Munmun De Choud- hury, and Diyi Yang. 2021. Latent hatred: A bench- mark for understanding implicit hate speech . InPro- ceedings of the 2021 Conference on Empirical Meth- ods in Natural Language Processing, pages 345ā 363, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics. Rui Fan, Shu Li, Tingting He, and Yu Liu. 2025.Aspect-based sentiment analysis with syntax- opinion-sentiment reasoning chain. InProceedings of the 31st International Conference on Computa- tional Linguistics, pages 3123ā3137, Abu Dhabi, UAE. Association for Computational Linguistics. Paula Fortuna and SĆ©rgio Nunes. 2018.A survey on au- tomatic detection of hate speech in text.ACM Com- put. Surv., 51(4). Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3309ā3326, Dublin, Ireland. Association for Computational Linguistics. Sirui Hong, Mingchen Zhuge, Jiaqi Chen, Yuheng Xiawu, and 1 others. 2024.Metagpt: Meta pro- gramming for a multi-agent collaborative frame- work.Preprint, arXiv:2308.00352. Cheng-Yu Hsieh, Chun-Liang Li, Chih-kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ran- jay Krishna, Chen-Yu Lee, and Tomas Pfister. 2023. Distilling step-by-step! outperforming larger lan- guage models with less training data and smaller model sizes. InFindings of the Association for Computational Linguistics: ACL 2023, pages 8003ā 8017, Toronto, Canada. Association for Computa- tional Linguistics. Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021.Lora: Low-rank adaptation of large language models.Preprint, arXiv:2106.09685. Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Ef- ficient memory management for large language model serving with pagedattention .Preprint, arXiv:2309.06180. Dianqing Lin, Tian Lan, Jiali Zhu, Jiang Li, Wei Chen, Xu Liu, Aruukhan, Xiangdong Su, Hongxu Hou, and Guanglai Gao. 2026. Exploring the capability boundaries of llms in mastering of chinese chouxi- ang language .Preprint, arXiv:2604.15841. Junyu Lu, Bo Xu, Xiaokun Zhang, Changrong Min, Liang Yang, and Hongfei Lin. 2023. Facilitating fine-grained detection of Chinese toxic language: Hierarchical taxonomy, resources, and benchmarks . InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16235ā16250, Toronto, Canada. Association for Computational Linguistics. Binny Mathew, Punyajoy Saha, Seid Muhie Yi- mam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021.Hatexplain: A benchmark dataset for explainable hate speech detection. InProceed- ings of the AAAI Conference on Artificial Intelli- gence, volume 35, pages 14867ā14875. Tarek Naous, Michael J Ryan, Alan Ritter, and Wei Xu. 2024.Having beer after prayer? measuring cultural bias in large language models. InProceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 16366ā16393, Bangkok, Thailand. Association for Computational Linguistics. Debora Nozza. 2021.Exposing the limits of zero- shot cross-lingual hate speech detection. InProceed- ings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Interna- tional Joint Conference on Natural Language Pro- cessing (Volume 2: Short Papers), pages 907ā914, Online. Association for Computational Linguistics. NicolĆ”s BenjamĆn Ocampo, Ekaterina Sviridova, Elena Cabrio, and Serena Villata. 2023. An in-depth anal- ysis of implicit and subtle hate speech messages. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Lin- guistics, pages 1997ā2013, Dubrovnik, Croatia. As- sociation for Computational Linguistics. OpenAI. 2026.Introducing GPT-5.4. Product an- nouncement. Nedjma Ousidhoum, Zizheng Lin, Hongming Zhang, Yangqiu Song, and Dit-Yan Yeung. 2019.Multi- lingual and multi-aspect hate speech analysis. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Lan- guage Processing (EMNLP-IJCNLP), pages 4675ā 4684, Hong Kong, China. Association for Computa- tional Linguistics. John Pavlopoulos, Jeffrey Sorensen, LĆ©o Laugier, and Ion Androutsopoulos. 2021.SemEval-2021 task 5: Toxic spans detection. InProceedings of the 15th International Workshop on Semantic Evalua- tion (SemEval-2021), pages 59ā69, Online. Associa- tion for Computational Linguistics. Paul Rƶttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert. 2021. HateCheck: Functional tests for hate speech detection models. InProceedings of the 59th Annual Meeting of the Association for Computational Lin- guistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 41ā58, Online. Association for Com- putational Linguistics. Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Ju- rafsky, Noah A. Smith, and Yejin Choi. 2020. So- cial bias frames: Reasoning about social and power implications of language. InProceedings of the 58th Annual Meeting of the Association for Compu- tational Linguistics, pages 5477ā5490, Online. As- sociation for Computational Linguistics. Anna Schmidt and Michael Wiegand. 2017.A survey on hate speech detection using natural language pro- cessing . In Proceedings of the Fifth International Workshop on Natural Language Processing for So- cial Media, pages 1ā10, Valencia, Spain. Associa- tion for Computational Linguistics. Kumar Shridhar, Alessandro Stolfo, and Mrinmaya Sachan. 2023.Distilling reasoning capabili- ties into smaller language models .Preprint, arXiv:2212.00193. Russell Spears. 2021.Social influence and group iden- tity. Annual Review of Psychology , 72:367ā390. Lei Wang, Min Huang, and Eduard Dragut. 2026. DanceHA: A Multi-Agent Framework for Document-Level Aspect-Based Sentiment Analysis. Proceedings of the AAAI Conference on Artificial Intelligence, 40(35):29714ā29722. Noah Wang, Z.y. Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, Man Zhang, Zhaoxiang Zhang, Wanli Ouyang, Ke Xu, Wen- hao Huang, Jie Fu, and Junran Peng. 2024. RoleLLM: Benchmarking, eliciting, and enhancing role-playing abilities of large language models. In Findings of the Association for Computational Lin- guistics: ACL 2024, pages 14743ā14777, Bangkok, Thailand. Association for Computational Linguis- tics. Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdh- ery, and Denny Zhou. 2023.Self-consistency im- proves chain of thought reasoning in language mod- els.Preprint, arXiv:2203.11171. Zeerak Waseem and Dirk Hovy. 2016.Hateful sym- bols or hateful people? predictive features for hate speech detection on Twitter. InProceedings of the NAACL Student Research Workshop, pages 88ā93, San Diego, California. Association for Computa- tional Linguistics. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023.Chain-of-thought prompting elicits reasoning in large language models .Preprint, arXiv:2201.11903. Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang. 2023. Autogen: Enabling next-gen llm ap- plications via multi-agent conversation.Preprint, arXiv:2308.08155. Yunze Xiao, Yujia Hu, Kenny Tsu Wei Choo, and Roy Ka-Wei Lee. 2024.ToxiCloakCN: Evaluating robustness of offensive language detection in Chi- nese with cloaking perturbations. InProceedings of the 2024 Conference on Empirical Methods in Natu- ral Language Processing, pages 6012ā6025, Miami, Florida, USA. Association for Computational Lin- guistics. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Day- iheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, and 41 others. 2025.Qwen3 technical report.Preprint, arXiv:2505.09388. Yongjin Yang, Joonkee Kim, Yujin Kim, Namgyu Ho, James Thorne, and Se-Young Yun. 2023.HARE: Explainable hate speech detection with step-by-step reasoning. InFindings of the Association for Com- putational Linguistics: EMNLP 2023, pages 5490ā 5505, Singapore. Association for Computational Linguistics. Marcos Zampieri, Skye Morgan, Kai North, Tharindu Ranasinghe, Austin Simmmons, Paridhi Khandel- wal, Sara Rosenthal, and Preslav Nakov. 2023. Target-based offensive language identification. In Proceedings of the 61st Annual Meeting of the As- sociation for Computational Linguistics (Volume 2: Short Papers), pages 762ā770, Toronto, Canada. As- sociation for Computational Linguistics. A Implementation and Artifact Details Table9gives the retainedQwen3-14Bconfigura- tion (Yang et al.,2025). Phase 1ā3 share the same local backbone, and Phase 4 is deterministic. The compact profile usesvLLM( Kwon et al.,2023). ItemValue BackboneQwen3-14B HardwareRTX 3090 + RTX 4090D Precisionbfloat16 Context length8192 Tensor parallelism 2 Phase 1local focus-unit segmentation; batch 4 Phase 23 roles; micro-batch 64; Chinese bystander with lexicon injection Phase 3quality threshold 0.60; cluster sim- ilarity 0.66; auditor enabled; max retained clusters 2 Runtime safeguards checkpoint resume; OOM split- retry Compact service profile vLLMbackend; auditor batch 1; re- generation disabled Phase 4deterministic reassembly Table 9: Main local configuration for the reported Qwen3-14B mainline. A.1 Runtime and Call Diagnostics Retained EN, ZH-main, and ZH-quadruple runs take 499.0, 617.6, and 623.3 minutes. Phase 2 makes 4,524, 6,978, and 6,933 role calls at 0.25ā 0.28 role-groups per second. Auditor refinement is triggered for 982, 1,645, and 1,610 groups. These figures describe retained runs; incomplete billing logs preclude a monetary cost estimate. A.2 Artifact Map and Recoverability TheartifactrootsSPAR/, SPAR_anonymous_submission/,and transfer_bundles/contain the determinis- tic TBO split and Phase 0 preparation path, Phase 1ā3 prompts and configurations, Phase 3 scoring and fallback code, direct and API baseline outputs, SynChain and Dance adaptations, diag- nostic summaries, and teacher/student exports. Historical bilingual distillation records are partly archival. Preserved reports and official evaluation exports support the reported construction and evaluation statistics, while the Chinese student bundle retains the independently recoverable training example. B Prompt Templates The three Phase 2 roles share the output con- tract, Target Resolution, Non-empty Argument, Output Sequence, Short Rationale, and Current Task blocks. Figure 4gives the Victim instruc- tion body and example headers. The Moderator and Cultural Bystander figures retain their role- specific additions. Repeated example titles are omitted. Phase 1 Segmenter # Role You are the SPAR-Hate Segment Agent specializing in English text. Your ONLY output is a strict JSON array. # Rules 1. Divide by target and intent: Split the text into separate groups if different targets are attacked or distinct intents exist. 2. Context isolation:group_textMUST be copied EXACTLY from the original text as a meaningful local context. Do not paraphrase, summarize, or alter punctuation/emojis. 3. Target normalization:canonical_targetis the direct entity mention. If the target is strictly absent, unmentioned, or null, output āIMPLICIT_TARGETā. 4. Output scope: ONLY outputlocal_group_id,group_text, and canonical_target. # Examples Example 1 (Multi-Target Segmentation) Example 2 (Null/Implicit Targets & General Venting) Example 3 (Non-harmful / Object Target with Emojis) # Current Task Original text:TEXT Output:[ Figure 3: Excerpt of the English Phase 1 segmenter prompt used in the bilingual main tracks. Victim # Role You are the āVictim Personaā Agent in a Hate Speech Evaluation system specializing in English text. You are highly sensitive, empathetic to marginalized groups, and acutely aware of emotional harm, exclusion, stereotyping, slurs, and microaggressions. Your ONLY output is a strict JSON array. # Rules 1. Victimās Lens: Put yourself in the shoes of the attacked group. If the text inside the focus area makes you feel unsafe, dehumanized, degraded, or stereotyped, flag it as harmful (harmful: 1). 2. Attention Lock (CRITICAL): Your evaluation area is STRICTLY limited to the text inside the<focus>and</focus>tags. You MUST NOT penalize offensive behaviors or slurs that exist outside these tags. 3. Target Resolution: prioritize the target inside<focus>; use outside context orTarget hintonly when the focus text is fragmentary; do not override an explicit in-focus target. If the focus only contains an insult or adjective but no addressee, output āIMPLICIT_TARGETā. 4. Argument Extraction:argumentMUST be an exact substring from INSIDE the<focus>tags and should be the shortest decisive span that carries the insult, stereotype, exclusion, or offensive action. 5. Multi-Argument Extraction (CRITICAL): Split independent attacks into separate objects. Do NOT merge two insults or predicates into one long argument just because they are joined by and, &, commas, or nearby context. 6. Non-empty Argument (CRITICAL):argumentmust NEVER be empty. Even for harmless outputs (harmful: 0), choose the shortest exact substring inside<focus>that best supports the harmless judgment. 7. Output Sequence: each object must strictly followrationale,target, argument,harmful. 8. Short Rationale:rationalemust be brief (under 30 words), single-line only, and must not contain raw newline characters inside the JSON string. # Examples Example 1 (Missing Subject Resolved via Context) Example 2 (Multiple Arguments on the Same Target within Focus) Example 3 (Multi-Argument Extraction on Permissible/Non-Harmful Text) # Current Task Target hint (soft only; use only when focus lacks a subject): "TARGET_HINT" Input:"FOCUS_TEXT" Output:[ Figure 4: Main instruction body of the English Victim prompt used in Phase 2. Moderator # Role You are the āModerator Personaā Agent in a Hate Speech Evaluation system specializing in English text. Your duty is to enforce social media platform Terms of Service (TOS) and community safety guidelines. You look for actionable offenses: severe slurs, targeted harassment, incitement of violence, and hate speech. You tolerate general profanity and subjective disagreements if they do not cross into targeted abuse. # Role-specific rules 1. Moderatorās Lens: Evaluate if the text INSIDE the focus area violates safety policies. Explicit slurs, hostile degradation, threats of violence, or targeted abuse must be flagged as harmful (harmful: 1). 4. Argument Extraction:argumentshould be the shortest decisive span that contains the actionable abuse, slur, threat, or hostile phrase. # Example headers Example 1 (Severe Violation with Pronoun Resolution) Example 2 (Explicit Slurs and Threats) Example 3 (Harmless/No Violation inside Focus) Figure 5: Role-specific English Moderator prompt used in Phase 2. Shared blocks with Victim are omitted. Cultural Bystander # Role You are the āCultural Bystander Personaā Agent in a Hate Speech Evaluation system specializing in English text. You act as an internet anthropologist. You focus on the sociological nuances of language: distinguishing reclaimed slang from derogatory usage, identifying group generalization templates, recognizing coded derogation, and differentiating between quoting/irony and genuine hostility. You assess the speakerās stance and tone. # Role-specific rules 1. Cultural Lens: Evaluate the societal and cultural implications of the text INSIDE the focus area. Recognize when a slur is reclaimed by the community versus when it is weaponized. Flag generalizations and coded derogation as harmful if they enforce negative societal stereotypes. Treat reclaimed slang as non-harmful only when the focus itself shows clear affiliative or celebratory usage. Negative complaints, commands, mockery, or group generalizations are not reclaimed evidence. 4. Argument Extraction:argumentshould be the shortest decisive span that captures the key cultural attack, slang, stereotype, or behavior. 5. Multi-Target/Argument Extraction (CRITICAL): Split independent attacks into separate objects. # Example headers Example 1 (Reclaimed Slang & Performative Tone) Example 2 (Group Generalization) Figure 6: Role-specific English Cultural Bystander prompt used in Phase 2. Shared blocks with Victim are omitted. Few-shot API Baseline You are a content review expert. Please identify hate speech in the sentence and extract one or more triplets. Your ONLY output is a strict JSON array. Each object must contain:rationale,target,argument,harmful. target: the discussed or attacked entity; if there is no explicit target, output āIMPLICIT_TARGETā. argument: the core evidence span. harmful: must be1or0. Output JSON array only. Do not output any explanation. Positive example: āThe men where I live are trash.ā Output:(target=men, argument=trash, harmful=1) Negative example: āWhat kind of dogs are the ugly gross ones with bloodshot eyes and long ass earsā Output:(target=dogs, argument=ugly gross ones with bloodshot eyes and long ass ears, harmful=0) Input:"TEXT" Output: return the answer starting exactly with[and ending with] Figure 7: English few-shot baseline prompt shared by the DeepSeek-v4-Pro and GPT-5.4 API baselines. Phase 3 Auditor ## Role You are the SPAR-Hate Phase-3 Arbitrate Auditor for English. You do not perform crude majority vote. You arbitrate between role evidence, candidate clusters, triage signals, and dynamic beacons to select the final tuple most likely to match the English ground truth. Your ONLY output is a strict JSON object. ## English Task Rules 1. Ground the final answer infocus_text.final_argumentmust be an exact or near-exact substring from the focus span. 2. Preferfinal_targetfrom inside the focus span. Only use canonical_targetor outside context when the focus span is fragmentary, pronominal, or missing the subject. 5. Distinguish reclaimed slang, quotation, refutation, irony, and performative speech from genuine targeted harm. 7. If one role misses coded derogation or over-penalizes a reclaimed or quoted expression, say so compactly infused_rationaleand correction_trace. 9. If several candidates differ mainly in argument length, prefer the shortest span that still independently expresses the attack core. ## Few-shot example headers Example 1: clear group generalization, final harmful Example 2: reclaimed slang, final non-harmful Example 3: shrink an over-broad argument span Figure 8: Rendered excerpt of the English Phase 3 au- ditor prompt used in the mainline configuration. C Baseline Adaptation SynChain and Dance are modified only enough to produce hate tuples compatible with the bench- marks. The aim is to preserve each methodās high- level inductive bias rather than rewrite it into a new multi-stage system ( Fan et al.,2025;Wang et al., 2026). General adaptation.Across all baselines, the adaptation layer makes only three minimal changes. First, raw outputs are normalised into strict JSON readable by the official evaluation scripts. Second, field names and label spaces are aligned: English usesharmful, Chinese uses hateful, andgroupis enabled only on the Chi- nese four-field track. Third, lightweight post- processing is applied only where required by the evaluation interface, including boundary cleaning, duplicate tuple removal, and unified sample-level export. None of these operations adds candidate arbitration or reranking. SynChain.The syntax-aware extraction bias is retained. The adapted workflow still generates structure-sensitive candidates first, then resolves target, argument, and label, and finally maps in- termediate candidates back to scorable text fields. Beyond alignment to the output contract, no extra multi-role generation, auditor arbitration, or lexi- con weighting is added. Dance.The divide-and-conquer skeleton of grouping, per-group inference, and merge is re- tained. The adapted system still performs in- ference over local groups before returning to sample-level outputs. The only change is that the internal extraction target is rewritten from aspect/opinion/sentiment-style fields into hate tu- ples. As with SynChain, no extra SPAR-style arbi- tration or lexicon-aware reranking is added. D Auditor Scoring and Bounded Arbitration Phase 3 scores, clusters, routes, and selects role candidates under grounding constraints. Table 10 gives the main settings, and Table11gives the scoring signals. ItemValue Role setvictim / moderator / cultural_bystander Quality threshold0.60 Cluster similarity0.66 Similarity weightstarget/argument/label/group= 0.45/0.45/0.05/0.05 Score bonusessupport role 0.12; candidate lexicon 0.05; clus- ter lexicon 0.08 Auditor lanesconsensus / defective / conflict / lexicon-hit / victim-microaggression Auditor batch2 (main 14B); 1 (compactvllm) Regenerationenabled (main 14B); disabled (compactvllm) Cluster retentionmax 2; score floor 1.05; min support roles 2 Decoding profilebfloat16; TP=2; context 8192 Final selectiondeterministic fallback after optional refinement Table 10: Key Phase 3 settings in the local 14B arbitra- tion configurations. Feature groupSignal Groundingargument in focus text; no invented evidence Target anchoringexplicit target; canonical-target match Structural validityJSON validity; legal label; track-compatible fields Group-label coherencegroupālabel consistency inZH-quadruple Soft priorslexicon bonus; moderator floor; victim mi- croaggression Table 11: Main feature groups used by the Phase 3 can- didate scorer. The quality threshold of 0.60 filters weak or un- grounded candidates before clustering. The simi- larity threshold of 0.66 limits merges between can- didates with divergenttargetorargumentfields. Both values are reused across the reported lan- guages and benchmarks. StepOperationBound ValidateCheck JSON, label, argument ground- ing, and target anchor qualityā„0.60 ClusterCompare target, argument, label, and optional group similarity ā„0.66 RouteAssign consensus, defective, or conflict lane one local unit RefineApply soft beacons and optional auditor call one regenera- tion SelectRank with deterministic fallbackat most two clus- ters Table 12: Operational summary of local Phase 3 arbi- tration. MeasureRetained-run summary Raw-cluster p95EN 3; ZH 4 Retained/final p95 and maximum 2 Conflict-lane share55.9ā61.7% Auditor-use share65.1ā70.7% Table 13: Boundedness and routing diagnostics across the retained runs. E Robustness and Semantic Validation E.1 Threshold Sensitivity The actual auditor is replayed on fixed seed-10947 subsets of 200 complete samples per language, covering 385 EN and 281 ZH focus groups per set- ting. Quality varies over0.55/0.60/0.65at simi- larity0.66. Similarity varies over0.60/0.66/0.72 at quality0.60. All other settings remain fixed. Ta- ble 14reports the retained summary. Language Observed range Default EN34.97ā35.0835.08 ZH26.28ā27.1226.82 Table 14: Four-hard-metric averages in the fixed-subset threshold replay. The results are not full-test signifi- cance estimates. E.2 Blind Semantic Audit Two independent bilingual or native-speaker re- viewers assess 60 stratified Chinese focus units with gold labels hidden. Separate final-tuple and Cultural-Bystander-rationale rubrics yield 120 judgements. Table 15gives acceptance and agree- ment. Consensus cases reach 70% strict overall acceptance; disagreement-heavy lexiconāauditor cases receive lower acceptance. E.3 Strict Sample Exact-Set Summary E.4 Attacked-Group Breakdown Table17reports ZH-quadruple scores by gold at- tacked group. The Racism+Sexism group has the largest TAH-to-Quad drop, locating the main loss in multi-group binding. The breakdown measures ObjectR1 R2 Agree.Īŗ Final tuple73.3 71.7 81.7 0.540 Bystander rationale 68.3 58.3 86.7 0.716 Table 15: Acceptance rates, percentage agreement, and CohenāsĪŗin the blind semantic audit. TrackSPAR-DeepSeekDirect references EN20/800 (2.50%)0ā9/800 (0.00ā1.13%) ZH-main 141/1605 (8.79%) zero-shot: 37/1605; 66/1605 Table 16: Sample exact-set summary. Some ZH few- shot baselines exceed the SPAR-DeepSeek value. robustness across the annotated groups; fairness requires separate evidence. Gold groupSamples TAH-S Quad-S Sexism302 36.66 34.49 Racism223 40.56 35.57 Region216 27.20 24.93 LGBTQ93 34.75 32.13 Racism + Sexism85 21.855.30 Table 17: ZH-quadruple attacked-group diagnostic. Scores are Soft F1 and groups are defined from gold labels. F Qualitative Cases and Error Analysis The Chinese case isolates a target conflict. The English case traces a conflict-lane decision within a multi-target sample. In this case, Victim and Cultural Bystander both identify the attacking target as āBeijingersā, whereas Moderator shifts the target to the men- tioned victim group. The auditor follows the grounded consensus on the focus text and down- weights the moderator reading. The current text is focus unitg 2 in a sample con- tainingg 1 ,g 2 , andg 3 . Phase 4 retains two harmful tuples fromg 1 ,(shit, shit, 1)fromg 2 , and one non-harmful tuple fromg 3 . One focus unit can emit several tuples. Forg 2 , the auditor selects the Victim reading and downweights the non-harmful Moderator and Cultural Bystander candidates. FieldContent TextThe people spreading the claim that Henan people steal manhole covers are in fact Beijingers. Gold tuple(Beijingers, saying that Henan people steal man- hole covers, 1) Victim(Beijingers, saying that Henan people steal man- hole covers, 1) Moderator(Henan people, steal manhole covers, 1) Cultural Bystander(Beijingers, saying that Henan people steal man- hole covers, 1) Triagelane=conflict; initial_conflict=target_conflict Beaconsgrounding;consensus Correction tracemoderator: downweighted for misreading the victim mention as the attacking target Final tuple(Beijingers, saying that Henan people steal man- hole covers, 1) Table 18: Chinese case study showing target-conflict resolution in Phase 3. FieldContent Textthat shit is not cute & dnt make u hard. Phase 4 restored sam- ple prediction (bitches, Ghetto, 1);(bitches, ignorant, 1); (shit, shit, 1);(IMPLICIT_TARGET, alldat shit, 0) Victim(shit, shit, 1) (u, dnt make u hard, 1) Moderator(IMPLICIT_TARGET, that shit, 0) (IMPLICIT_TARGET, dnt make u hard, 0) Cultural Bystander(shit, shit, 0) Triagelane=conflict Beaconsgrounding,victim_microaggression Correction tracemoderator:downweighted; cultural_bystander: downweighted Final tuple(shit, shit, 1) Table 19: English case study showing conflict-lane ar- bitration inside a multi-target sample. G Distillation Data and Student Training G.1 Teacher-trace Schema The distillation target includes more than the final tuple. Each training row also preserves intermedi- ate decomposition, role candidates, conflict diag- nosis, and the final arbitration result, so that the student learns a structured decision process rather than only the flat output surface. G.2 Filtering and Dataset Statistics The bilingual statistics come from the main dis- tillation construction run. The pipeline has three steps: first, retain train rows with recoverable iden- tifiers and serialisable teacher packages; second, build a strict HQ reasoning subset using ground- edness, final tuple validity, and audit confidence; third, expand it under the same grounding con- straints into reasoning-plus and JSON-only vari- ants for curriculum training. G.3 A Real Constructed Training Example Table 22shows one real final training row from the released Chinese student bundle. The sample Field groupContent Identifierssample_id,group_id, source, split, language Source textoriginal text and the current fo- cus/group unit Focus evidencegroup_text,focus_text, canonical_target Role evidencetuples from the three roles, short rationales, quality scores, lexicon hit Conflict diagnosistriage lane, field conflicts, applied beacons, clusters Teacher decisionteacher_final_tuple, teacher_rationale,audit confidence Supervision messages system instruction, user evidence package, assistant supervision tar- get Table 20: Schema of the final constructed distillation rows. StageCountNote Phase 3 group traces 7,752all available local decision traces Packaged train rows3,739base samples entering the teacher package Teacher reasoning HQ2,221main warm-up variant HQ Chinese rows1,590zh_group_hateful HQ English rows631en_span_harmful Teacher reasoning plus3,194reasoning variant with re- laxed conflict threshold Teacher JSON-only plus 3,194final JSON-aligned variant Validation checksall passduplicate id 0; schema error 0; invalid final tuple 0 Table 21: Summary statistics for the teacher-trace con- struction pipeline. comes from a multi-target Chinese input, but the current training row supervises only the second fo- cus unit. The three role rationales are retained so that the student learns argument anchoring when targets agree but arguments compete. G.4 Student Training Setup ItemValue Student backboneQwen3-4B-Instruct-2507 Prompt templateQwen3 no-think Training splittrain only Supervision regime teacher-only Curriculum stagesteacher reasoning HQ; teacher reasoning plus; teacher JSON- only plus Epochs per stage1.0 / 1.0 / 1.0 Sequence length4096 Per-device batch size 1 Gradient accumula- tion 8 Optimizer and sched- ule learning rate 1e-4; cosine; warmup ratio 0.05 Precisionbf16 Adaptation method QLoRA (4-bit BnB) LoRA settingrank 32; alpha 32; dropout 0.0 FieldContent Source textMarriage is a necessity for women. Men can stay unmarried for life, but women cannot. In reality, countless men are even more eager. They will empty the family savings and act like simps just to marry and have children. Absurd. Sample-level gold(marriage, a necessity for women, 1); (countless men, empty the family savings and act like simps just to marry and have children, 1) Current focusIn reality, countless men are even more eager. They will empty the family savings and act like simps just to marry and have children. Absurd. Dataset variantteacher_trace_reasoning_hq Victimtuple: (men, even more eager than women, and willing to empty the family savings and act like simps just to marry and have children, harmful); rationale: It describes men as simps and derides their behaviour, carrying clear gender stereotyping and disparagement. Moderatortuple: (men, willing to empty the family savings and act like simps just to marry and have chil- dren, harmful); rationale: The focus uses deroga- tory wording such as āsimpsā to negatively charac- terise men. The phrasing is attacking and demean- ing, which violates community standards. Cultural Bystandertuple: (men, act like simps, harmful); rationale: [Lexicon Hit] āsimpā is used here to belittle men and suggest a lack of dignity. Conflict diagnosislane=consensus; high-risk field=argument; error tags=lexicon_hit,auditor_span_refined Teacher final tupletarget=men;argument=act like simps; group=unknown; harmful=1; direction=harmful Assistant targetformat:reasoning+JSON; final JSON: target=men;argument=act like simps; group=unknown; harmful=1; direction=harmful Table 22: A condensed real training row from the con- structed distillation data. Table 23:Student-model training setup used in the dis- tillation pipeline. Student training follows a teacher-only curricu- lum. Stage 1 uses the strict HQ reasoning subset to establish the structured decision process. Stage 2 expands coverage with reasoning-plus. Stage 3 aligns the model to strict JSON output. Under this protocol, training uses only the train split, ground truth is never used as assistant supervi- sion, evaluation is limited to teacher reasoning and teacher JSON, and manual prefill is disal- lowed. TheQwen3-4Bbackbone, 4-bit QLoRA, and LoRA rank/alpha settings follow the corre- sponding parameter-efficient fine-tuning methods ( Yang et al.,2025;Dettmers et al.,2023;Hu et al., 2021). H Full Bilingual Ablation Results Table24gives the complete bilingual ablations underlying the compact analysis in Section4.4. (a)ZH-main SettingTargetArgumentT-A PairT-A-H Tri. Avg. Hard Soft Hard Soft Hard Soft Hard Soft Qwen3-14B (mainline) 48.73 55.75 21.03 56.85 12.36 34.14 8.71 24.97 32.82 w/o phase1 segmenter 52.34 64.46 15.20 48.92 9.26 34.43 7.02 26.39 32.25 w/o phase2 victim50.40 60.67 15.95 46.48 10.53 32.84 7.63 24.78 31.16 w/o phase2 moderator 49.90 60.07 16.51 47.65 10.86 34.91 7.75 26.49 31.77 w/o phase2 bystander 50.18 61.11 16.10 50.90 10.50 36.93 8.20 28.90 32.85 w/o phase346.54 56.64 17.03 51.42 10.67 36.58 7.68 27.10 31.71 (b)EN-main SettingTarget Arg. Targeted NTA TA Harm Avg. Qwen3-14B (mainline) 43.45 47.96 34.68 20.96 18.68 59.31 37.51 w/o phase1 segmenter 46.95 47.11 34.06 20.65 18.04 57.21 37.34 w/o phase2 victim44.70 49.17 32.68 21.18 17.20 57.21 37.02 w/o phase2 moderator 44.29 48.25 34.07 21.34 17.51 59.11 37.43 w/o phase2 bystander 44.49 47.24 33.10 20.64 18.37 59.22 37.18 w/o phase343.66 46.77 34.07 20.42 17.81 57.61 36.72 Table 24: Complete bilingual Qwen3-14B ablations. I Full Results on ZH-quadruple Table25reports the full four-field results used for the Chinese extension discussed in Section4.3. ModelTargetArgumentT-A PairT-A-H Tri.Quad.Avg. Hard Soft Hard Soft Hard Soft Hard Soft Hard Soft Local LLMs 0-shot-Qwen3-14B43.94 53.81 15.60 47.00 9.92 29.58 7.65 23.18 6.87 19.42 25.70 SPAR-Qwen3-14B (mainline) 48.5458.5819.61 53.9411.0034.148.5126.616.18 22.3728.95 SPAR-Qwen3.5-35B50.81 60.8419.5752.0513.64 38.50 10.59 29.94 8.68 24.48 30.91 API Models LLaMA3-70B*30.87 41.45 14.80 46.68 8.29 25.38 7.40 22.58 4.72 13.14 21.53 Claude-3.5-Sonnet*41.45 54.06 15.80 55.80 10.43 36.96 9.28 33.047.10 25.22 28.91 few-shot-DeepSeek-v4-Pro54.86 65.7018.62 60.50 14.39 41.02 11.34 32.51 9.7027.9933.66 few-shot-GPT-5.452.8764.9817.09 60.28 12.43 42.17 9.45 31.54 8.15 27.50 32.65 SPAR-DeepSeek-v4-Pro52.66 61.2923.73 65.66 16.4543.12 12.7032.25 10.8227.6534.63 SPAR-GPT-5.452.15 60.63 21.6955.34 16.3743.71 13.15 33.40 11.63 29.1133.72 Table 25: Full Chinese four-field results. Starred rows are historical STATE-ToxiCN values (Bai et al.,2025).