Paper deep dive
VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation
Mingyu Yuan, Shengtao Wen, Lingbing Guo, Zhen Bi, Xiang Chen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/22/2026, 2:36:40 AM
Summary
The paper introduces VARM-Bench, a benchmark for evaluating verifiable structured reasoning in Chinese abusive speech moderation. It addresses the limitation of existing benchmarks that rely on final labels by providing field-anchored chain-of-thought rationales. The benchmark includes 8,000 instances with six explicit decision fields: target, target type, target explicitness, author stance, harmfulness label, and fine-grained category. A deterministic protocol evaluates these fields without LLM judges, revealing that strong label-level performance can conceal significant errors in complete moderation records.
Entities (16)
Relation Signals (15)
VARM-Bench → evaluatesmodel → Qwen3.7-Max
confidence 95% · we evaluate language models across multiple model families... Qwen3.7-Max
VARM-Bench → evaluatesmodel → DeepSeek V4 Pro
confidence 95% · we evaluate language models across multiple model families... DeepSeek-V4-Pro
VARM-Bench → evaluatesmodel → Qwen2.5-7B
confidence 95% · we evaluate language models across multiple model families... Qwen2.5-7B
VARM-Bench → evaluatesmodel → LLaMA-3.1-8B
confidence 95% · we evaluate language models across multiple model families... Llama-3.1-8B
VARM-Bench → evaluatesmodel → InternLM3-8B
confidence 95% · we evaluate language models across multiple model families... InternLM3-8B
VARM-Bench → evaluatesmodel → GPT-5.5
confidence 95% · we evaluate language models across multiple model families... GPT-5.5
VARM-Bench → supportstask → Chinese Abusive Speech Moderation
confidence 95% · We introduce VARM-Bench, a benchmark for field-anchored chain-of-thought rationales in Chinese abusive-speech moderation.
Lingbing Guo → affiliatedwith → Nanjing University
confidence 90% · Lingbing Guo 2 ... 2 Nanjing University
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text. Existing Chinese benchmarks support label classification, fine-grained toxicity categorization, and target-aware extraction, but do not provide a unified representation for deterministically verifying the stated basis of a moderation decision. We introduce VARM-Bench, a benchmark for field-anchored chain-of-thought rationales in Chinese abusive-speech moderation. Each instance contains a concise natural-language rationale with explicit anchors for six decisions: target, target type, target explicitness, author stance, harmfulness label, and fine-grained category. Our deterministic protocol evaluates field correctness, target alignment, output validity, complete-record agreement, and hidden record errors conditioned on correct final decisions, without relying on an LLM judge. Under a common structured-output protocol, we evaluate language models across multiple model families using zero-shot prompting, taxonomy guidance, and structured CoT supervision, and analyze lexical-cue sensitivity and field-level errors. Results show that strong label-level performance can conceal substantial errors in complete moderation records. VARM-Bench provides an auditable and reproducible benchmark for evaluating verifiable moderation rationales in Chinese abusive-speech moderation.
Tags
Links
- Source: https://arxiv.org/abs/2608.15600v1
- Canonical: https://arxiv.org/abs/2608.15600v1
Trouble viewing inline? Open PDF directly →
Full Text
46,953 characters extracted from source content.
Expand or collapse full text
VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation Mingyu Yuan 1∗ , Shengtao Wen 1∗ , Lingbing Guo 2 , Zhen Bi 3 , Xiang Chen 1† 1 MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics 2 Nanjing University 3 College of Computer Science, Huzhou Normal University xiang_chen@nuaa.edu.cn Abstract The widespread circulation of abusive online content has in- creased the need for reliable moderation of Chinese social- media text. Existing Chinese benchmarks support label classi- fication, fine-grained toxicity categorization, and target-aware extraction, but do not provide a unified representation for de- terministically verifying the stated basis of a moderation de- cision. We introduce VARM-Bench, a benchmark for field- anchored chain-of-thought rationales in Chinese abusive- speech moderation. Each instance contains a concise natural- language rationale with explicit anchors for six decisions: tar- get, target type, target explicitness, author stance, harmfulness label, and fine-grained category. Our deterministic protocol evaluates field correctness, target alignment, output validity, complete-record agreement, and hidden record errors condi- tioned on correct final decisions, without relying on an LLM judge. Under a common structured-output protocol, we eval- uate language models across multiple model families using zero-shot prompting, taxonomy guidance, and structured CoT supervision, and analyze lexical-cue sensitivity and field-level errors. Results show that strong label-level performance can conceal substantial errors in complete moderation records. VARM-Bench provides an auditable and reproducible bench- mark for evaluating verifiable moderation rationales in Chi- nese abusive-speech moderation. Code — https://github.com/NUAA-MMMI/VARM-Bench Dateset — https://huggingface.co/datasets/NUAA- MMMI/VARM-Bench Content warning: This paper includes examples containing language that some readers may find offensive or vulgar. Introduction Abusive-language detection is commonly evaluated using final-label accuracy (Founta et al. 2018). A correct label, however, can conceal an incorrect basis for the moderation decision. A model may identify the wrong referent, mis- take quoted or rejected abuse as the author’s own position, ∗ These authors contributed equally. † Corresponding author. or assign a harm category that is inconsistent with the ac- tual target and discourse context while still producing the correct final label. These errors are particularly consequen- tial in Chinese social-media text, where implicit references, homophonic substitutions, negation, quotation, and sarcasm can change who is being discussed, and whether apparently offensive wording is abusive in context. Existing Chinese benchmarks have advanced from offensive-language classification and social-bias detection to implicit-toxicity analysis, fine-grained categorization, and target-aware structured extraction (Deng et al. 2022; Jiang et al. 2022). Nevertheless, labels, categories, and target spans represent only isolated components of a moderation decision. A target span does not reveal whether the author attacks, re- jects, quotes, or neutrally mentions the target. Conversely, a structured field tuple does not explain which input cues sup- port its predictions. Free-form rationales can express these relations, but their central claims are difficult to extract and evaluate consistently. As shown in Figure 1, reliable eval- uation therefore requires a unified representation in which the stated reasoning is expressed in natural language, its key moderation decisions are explicitly recoverable, and the re- sulting record can be checked deterministically. We introduce VARM-Bench (Verifiable Abusive Reason- ing Moderation Benchmark), for field-anchored chain-of- thought (CoT) in Chinese abusive-speech moderation. For each input, a model generates one explanation with explicit anchors for six decisions: target, target type, target explicit- ness, author stance, harmfulness, and fine-grained category. A deterministic parser converts these anchors into a complete moderation record, making every scored decision traceable to the explanation and verifiable against the input. We treat the generated CoT only as an inspectable rationale, not as a faithful account of the model’s internal reasoning. VARM-Bench contains 8,000 Chinese social-media com- ments, including 1,440 challenging non-harmful examples involving quotation, opposition, neutral identity mentions, behavioral criticism, and other misleading surface cues. Af- ter model-assisted pre-annotation, every reference record was manually corrected and rechecked against a shared codebook. Evaluation covers final-label and category performance, as well as deterministic agreement between records extracted arXiv:2608.15600v1 [cs.AI] 16 Aug 2026 What is actually being judged? Can keywords determine harmfulness? Does a correct label mean correct reasoning? Current models rely on trigger words instead of understanding context. label: Harmful Reason: direct personal insult label: Harmful Reason: refuting insult, not endorsing it Same keyword ≠ Harmfulness 这算法很糟糕。 This algorithm is terrible. 你就是个废物,什么事都做不好,还总拖别人后腿! You are a worthless loser; you do everything badly and always hold others back. 别再张口闭口骂别人“废物”了,批评问题可以,但没必要人 身攻击! Stop calling others “ worthless losers ” ; criticize the problem, not the person. X人都很蠢。 People from X are stupid. 你真是个废物。 You are a worthless. Target: Work / Issue Label: Non-harmful Target: Group Label: Harmful Target: Person Label: Harmful Models must identify the target of evaluation rather than the sentiment itself. Shortcut Path (Spurious) Semantic Path (Ideal) Same prediction, different reasoning. Label-only evaluation cannot tell the difference. matched keyword (“loser/废物”) Trigger word detected Predict harmful Identify target (e.g., Person) Understand intention (personal attack) Predict harmful Final Label (e.g., harmful) Input comment Figure 1: Motivation for VARM-Bench. Similar lexical cues can correspond to attacks, quotations, opposition, or neutral mentions. A final label alone cannot reveal whether the target, stance, and harm category are jointly correct. from output CoTs and their references. Separate audits assess sensitivity to reference conventions and explanation quality. We establish a unified protocol for generating, recovering, and evaluating complete moderation records. We use this pro- tocol to benchmark multiple open and closed models across prompting and supervision settings and provide diagnostic analyses of hidden record errors, context-dependent lexical cues, field-level bottlenecks, and generated-rationale quality. Together, the benchmark, protocol, and analyses provide a reproducible foundation for studying verifiable reasoning in Chinese abusive-speech moderation. Our main contributions are summarized as follows: • We introduce a Chinese abusive-speech benchmark in which six moderation decisions are embedded in one natural-language CoT and deterministically reconstructed as a complete moderation record. • We establish a multi-level evaluation protocol separating decision quality, field correctness, record validity, and complete-record agreement, with human review and au- dits of references and rationales. • We evaluate open and closed models on VARM-Bench across prompting and supervision settings and analyze hidden record errors, lexical-cue sensitivity, and field bot- tlenecks, identifying referent localization as the main bot- tleneck under frozen-reference scoring. Related Work Abusive-Language Detection and Chinese Benchmarks. Early abusive-language datasets mainly evaluated post-level classification (Davidson et al. 2017). Later work broadened evaluation to implicit toxicity and target-based offensive- language identification (Hartvigsen et al. 2022; Zampieri et al. 2023). Toxic-span benchmarks and functional tests fur- ther examine localized evidence and behavioral failures in hate-speech systems (Pavlopoulos et al. 2021; Röttger et al. 2021). Chinese benchmarks make parallel progress. COLD and SWSR cover offensive language and online sexism (Deng et al. 2022; Jiang et al. 2022); CDial-Bias and ToxiCN add targeted groups, implied attitudes, toxicity types, and expres- sion forms (Zhou et al. 2022; Lu et al. 2023). STATE-ToxiCN moves to span-level Target–Argument–Hateful–Group ex- traction, while ChineseHarm-Bench covers policy-violation categories with an expert-curated rule base (Bai et al. 2025; Liu et al. 2025). Contextual resources such as Social Bias Frames and Latent Hatred show why intent, social implica- tion, and implicit-hate explanations matter beyond surface toxicity (Sap et al. 2020; ElSherief et al. 2021). Moderation studies further show that decisions vary with annotation def- initions and conversational context (Schöpke-Gonzalez et al. 2023; Yu, Blanco, and Hong 2022). Policy goals and modera- tor needs also affect how similar cues should be handled (Cao et al. 2023; Zheng et al. 2024). This work motivates evalu- ation beyond a final label or isolated span: the target, its properties, author stance, harmfulness, and category must be checked together to determine whether a moderation deci- sion rests on the right basis. Existing benchmarks provide labels, taxonomies, target annotations, and structured extrac- tion, but none reconstructs this complete moderation record from one generated rationale. Rationales and Structured Verification. Rationale benchmarks such as HateXplain and ERASER provide human-annotated evidence for model predictions (Mathew et al. 2021; DeYoung et al. 2020). Such evidence makes predictions easier to inspect, but extracted spans do not show how the target, author stance, harmfulness label, and Multi-Source Comment Collection ... Topics: Gender, Ethnicity, Region, LGBTQ+, Disability, Status ... Filtering Normalize formating Anoymize mentions Remove emojis ... Candidate Comments #Text#: <The input text to be evaluated> #Target#: <What is the target?>, <What type?>, <Explicit or implicit?> #Stance#: What is the stance? (A) attack_or_endorse (B) oppose_attack (C) ... #Label#: Is it harmful? (A) harmful (B) non-harmful #Category#: Which harm category? (A) regional (B) gender (C) ... #CoT#: Explain the choices above in natural language. Annotation Rules Source-Text Grounding. Complete-Record Coverage. Rationale Sufficiency. Structured CoT Annotation Construction LLM Pre-annotationHuman Review Deterministic Evaluation of Structured Moderation Records Structured CoT Annotation Construction Anchor Parser Target Identifacation Scope & Explicitness Stance Interpretation Harm Classification Record-Level Evalution "text": "小县城公务员也就靠个编制,离了 编制啥也不是.", “ gloss": "Small-town civil servants just rely on their government tenure; without it, they are nothing.", " target": "Small-town civil servants", " target_type": "group", " target_explicitness": "explicit", " stance": "attack_or_endorse", " label": "harmful", " category": “status/occupation attack ", The comment explicitly targets[T: small-town civil servants] with the claim that they are “nothing” without secure public-sector status. The target denotes an occupational group [TY: group] that is overtly named[T: explicit]. By denying the group’s ability and worth, the author adopts [S: attack_or_endorse]. This generalized derogation is [L: harmful]; because it rests on occupational identity and institutional status rather than regional origin, it falls under [C: status/occupation]. Figure 2: Overview of the data construction, annotation, and validation pipeline. Multi-source Chinese comments are normalized, anonymized, and filtered, annotated through LLM-assisted pre-annotation and human review, and finally evaluated at the target, scope, stance, harm, and record levels. fine-grained category jointly support a moderation decision. Chain-of-thought prompting made natural-language reason- ing a common prediction interface (Wei et al. 2022). Self- consistency showed that sampling multiple reasoning paths can improve prediction reliability (Wang et al. 2023b). Later work evaluated intermediate reasoning with chain-level met- rics and process supervision (Golovneva et al. 2023; Prasad et al. 2023). Stepwise verification further formalized process- level supervision for complex reasoning (Lightman et al. 2024). Model-based evaluators extended automatic assess- ment to broader judgment tasks (Zheng et al. 2023). Safety benchmarks test complementary properties such as safety knowledge, trustworthiness, adversarial robustness, and re- fusal behavior (Zhang et al. 2024; Wang et al. 2023a). Harm- Bench focuses more directly on automated red teaming and harmful-behavior evaluation (Mazeika et al. 2024). Natural- language rationales can express relations among modera- tion fields, but their core judgments cannot be extracted and compared consistently without explicit checkpoints. Gener- ated CoT explanations may also diverge from the factors behind predictions (Lanham et al. 2023; Turpin et al. 2023). In VARM-Bench, each natural-language CoT contains six checkpoints. A deterministic parser reconstructs the com- plete record, and a manual audit evaluates grounding, con- textual correctness, and support for the predicted decision. Preliminaries Motivation: Why Verifiable Structured Reasoning? Final-label evaluation reveals a moderation decision but not whether that decision follows from a correct interpretation of the input. A model may return the correct label while identifying the wrong target, misinterpreting target explic- itness or author stance, or assigning an incompatible cat- egory. Field tuples support deterministic scoring but omit input-specific justification. Free-form rationales capture con- textual relations, but they are difficult to extract and com- pare consistently. VARM-Bench combines these functions in a single natural-language moderation CoT with six ex- plicit anchors for target, target type, target explicitness, au- thor stance, harmfulness label, and fine-grained category. A deterministic parser reconstructs the stated record for field- level and joint scoring, while the surrounding rationale en- ables audits of grounding, contextual correctness, and infer- ential sufficiency. In this paper, verifiable reasoning refers only to this operational property. It does not imply access to latent model computation. Task Definition Each input x i is paired with a six-field reference moderation record z i = (t i ,τ i ,ε i ,s i ,y i ,c i ),(1) and a field-anchored CoT r i that justifies its decisions. The benchmark is D =(x i ,z i ,r i ) N i=1 .(2) Given only the input text x i , the model generates a single natural-language moderation CoT: ˆr i = f θ (x i ).(3) The generated CoT uses this fixed sequence of six anchors: [T:]→ [TY:]→ [T:]→ [S:]→ [L:]→ [C:], which encode the target, target type, target explicitness, au- thor stance, harmfulness label, and fine-grained category, respectively. A deterministic parserP extracts the anchored values and reconstructs the predicted moderation record: ˆz i =P(ˆr i ) = ( ˆ t i , ˆτ i , ˆε i , ˆs i , ˆy i , ˆc i ).(4) No separate field tuple is generated: all field-level predictions are parsed directly from the same generated CoT, whose surrounding text is retained for rationale auditing. The six parsed fields, together with their operational definitions and admissible values, are specified as follows: • Target (t i ) identifies the shortest stable referent of the main decision-relevant proposition. Incidental entities are excluded, while coordinated referents under the same proposition are treated as one target. In quoted or op- posed abuse, the target remains the referent of the embed- ded claim. Output: a normalized free-text referent or no clear target. • Target Type (τ i ) describes the referential scope of the target. Values: single object, group object, and no clear target. • Target Explicitness (ε i ) indicates whether the target is directly expressed or inferred from context. Values: ex- plicit, implicit, and no clear target. • Stance (s i ) describes the author’s relation to the relevant abusive claim. Values: attack or endorse, oppose attack, quote or report, and neutral mention. • Label (y i ) indicates whether the text performs or endorses abuse. Values: harmful and non-harmful. • Category (c i ) identifies the primary basis of the mod- eration decision. Values: general abuse, region/ethnicity, gender, sexual orientation and gender identity (SOGI), status/occupation, body/health, and non-attack. All non- harmful instances use non-attack. Detailed tie-breaking rules appear in the Appendix. Prose between anchors must link input-specific evidence to deci- sions, especially for context-sensitive phenomena such as quotation, negation or opposition, implicit reference, sar- casm, and behavioral criticism. Listing or paraphrasing an- chor values may satisfy the format but not the rationale re- quirement. Automatic evaluation measures field-level and joint agreement with the reference record. Benchmark Construction Figure 2 summarizes the construction pipeline. VARM- Bench contains 8,000 records split into 5,600 training, 800 development, and 1,600 test instances, with a consistent 55/45 harmful–non-harmful ratio and 1,440 difficult non- harmful cases. To prevent cross-split leakage, normalized duplicates and pairs with RapidFuzz similarity of at least 95% are confined to a single split. Data Collection and Filtering We collected public posts and comments from Bilibili, Zhihu, Baidu Tieba, and Hupu. Topic-based queries provided can- didates for the six harmful categories. Related lexical cues TrainTestDev 01000200030004000 Non-attack 3,600 General abuse 733 Origin/culture 733 Gender 733 SOGI 733 Status/role 734 Body/health 734 Number of instances Figure 3: Seven-way category distribution of VARM-Bench across the train, development, and test splits. were also used to find difficult non-harmful examples, in- cluding quotations of abusive language, objections to abuse, neutral identity references, personal accounts, and criticism of behavior. These examples were labeled as non-harmful and assigned to the non-attack category. Search terms were used only for retrieval; annotators made the final decisions from the referent, the author’s stance, and the full context. We kept self-contained texts of at least 15 characters and discarded advertisements, corrupted text, meaningless sym- bol strings, and replies that required missing context. We re- moved URLs and markup and deleted or generalized personal identifiers without changing the intended meaning. Finally, we applied Unicode NFKC normalization, removed exact du- plicates, and excluded text pairs with RapidFuzz similarity of 95% or higher, including cross-split pairs. Annotation Schema and Procedure LLM-assisted pre-annotation produced a draft six-field record and an 80–180-character anchored CoT. Three mas- ter’s students and one doctoral student then independently reviewed all 8,000 input–record–rationale triples using a shared codebook. The first pass revised every field and ex- planatory link; the second verified each record against the source text and flagged unresolved cases. Accepted ratio- nales contained each anchor once, in order. The codebook specified tie-breaking rules for category boundaries, implicit targets, quoted or opposed abuse, and multi-attribute cases. Quoted, opposed, or neutral mentions remained non-harmful unless the author endorsed the attack, while the attribute supporting the main abusive claim determined the primary category. Reviewers revised or rejected unsupported targets, incorrect stance attributions, incompatible label–category as- signments, anchor-only restatements, and unsupported input- to-decision links. The final release retained only source- grounded, mutually consistent, human-verified records. Annotation Quality Control. A record was accepted only when its target attributes, stance, label, category, and rationale were supported by the source text and mutually consistent. A dataset-wide valida- tor checked completeness, valid values, anchor order and ModelSettingDecisionStructured FieldsValidityJointHidden Error Label Cat. T-F1 Type Exp. Stance Parse Format JREM HER-C HER-L Qwen2.5-7BZero-shot88.4 63.7 57.5 50.1 34.954.494.091.631.954.962.8 Qwen2.5-7B+Taxonomy 90.6 76.7 58.0 48.9 36.853.297.194.036.454.859.3 Qwen2.5-7BCoT-SFT96.7 88.2 73.1 59.2 46.978.499.699.459.534.438.3 Llama-3.1-8BZero-shot61.4 54.8 50.0 51.0 34.130.686.682.124.648.061.6 Llama-3.1-8B+Taxonomy 80.1 63.6 52.2 53.9 34.644.394.291.330.451.761.7 Llama-3.1-8BCoT-SFT95.3 86.3 69.2 61.5 46.080.999.599.356.037.241.1 InternLM3-8BZero-shot77.6 56.7 31.8 44.4 24.342.577.853.37.986.388.4 InternLM3-8B+Taxonomy 79.9 63.9 36.7 47.1 25.852.083.472.012.779.882.8 InternLM3-8BCoT-SFT96.1 88.1 72.3 65.7 53.775.4 100.099.859.534.538.1 GPT-5.5Zero-shot97.6 85.7 69.1 59.5 39.276.1 100.0 100.055.438.143.2 Qwen3.7-MaxZero-shot96.9 86.4 61.7 60.7 40.666.699.999.449.645.048.9 DeepSeek-V4-Pro Zero-shot95.6 86.8 61.1 63.0 44.067.099.097.446.348.151.4 Table 1: Performance comparison of different models on the deduplicated 1,600-item test set (%) across classification, target extraction, structural validity, and complete-record evaluation. Classification metrics use Macro-F1, while T-F1 uses mean character-F1. HER-C/L are lower-is-better; all other metrics are higher-is-better. Bold indicates the best result in each column. uniqueness, and cross-field constraints. Failed records were returned for human correction. All 8,000 released records passed these checks, and no cross-split duplicates remained at the 95% RapidFuzz threshold. To assess annotation con- sistency, three trained reviewers independently reannotated a category- and phenomenon-stratified sample of 200 records. Reviewers saw only the source text and had no access to the frozen reference or each other’s decisions. Krippendorff’s α was 0.943 for Label, 0.935 for Category, 0.914 for Stance, 0.797 for Target Type, and 0.671 for Target Explicitness. Target is a normalized free-text referent, so its agreement was measured with pairwise character-F1, yielding 0.775. These results show high agreement on core moderation deci- sions, with greater contextual sensitivity in Target Explicit- ness. Disagreements were adjudicated after all independent judgments were recorded. Evaluation Metrics Metrics are computed end to end: unparseable outputs remain in the denominators and receive zero. We report Macro-F1 for the five discrete fields, normalized character-overlap F1 (T-F1) for the free-text target, and Parse Success when all six valid anchors occur once and in order. JREM follows the frozen reference convention rather than asserting a uniquely correct semantic record. Let J i = 1 when output i is parseable, its target T-F1 is at least 0.5, and the other five fields match the reference; otherwise J i = 0. Let M X i indicate a correct category (X = C) or label (X = L) prediction: JREM = 1 N N X i=1 J i ,(5) HER-X = P N i=1 M X i (1− J i ) P N i=1 M X i , X ∈C,L. (6) HER-C and HER-L therefore measure hidden complete- record errors among category- and label-correct outputs; parse failures lower JREM but do not enter their denomi- nators. Higher is better except for HER-C and HER-L. Experiments Experimental Setup We evaluate Qwen2.5-7B, Llama-3.1-8B, and InternLM3- 8B under zero-shot, taxonomy-guided zero-shot, and CoT- SFT; GPT-5.5, Qwen3.7-Max, and DeepSeek-V4-Pro use zero-shot. All systems share the deduplicated 1,600-item test set. CoT-SFT uses 5,600 training and 800 development in- stances with prompt-masked LoRA and development-loss early stopping, directly supervising all six anchors and their connecting prose. Across settings, the test inputs, six-anchor output contract, target normalization, parser, and metrics remain fixed; only the prompting or adaptation condition changes. Each table entry is recomputed from its frozen full- test output, with no manual repair of parse or field errors. The test set is excluded from adaptation and model selection. De- terministic decoding and one parser score every output; the Appendix documents prompts, hyperparameters, hardware, and implementation needed for reproduction. Main Results Q1: How often do correct labels conceal record errors? To assess whether label-level performance reflects complete- record quality, we evaluate all models and training settings on the same 1,600-item test set using JREM and HER- C/HER-L. JREM requires the target to meet the character- overlap threshold and the other five fields to match the refer- ence. HER-C/HER-L measure the record errors that remain among parseable outputs with a correct category or label. Table 1 shows that GPT-5.5 achieves Label and Category Macro-F1 scores of 97.6% and 85.7%, but its JREM is only 55.4%; 38.1% of its category-correct outputs still contain a record error. For the three open models, taxonomy guid- ance raises Category Macro-F1 by 7.2–13.0 points, while 0 20 40 60 80 100 JREM (%) (a) Demographic identity 74.2 57.5 70.8 61.2 73.3 48.8 80.8 46.2 67.5 43.8 64.2 36.2 (b) Social status/role 71.4 50.0 71.4 54.8 79.4 52.4 74.6 47.6 77.8 40.5 69.8 35.7 (c) Body/health/disability 54.4 42.1 53.5 44.7 60.5 35.5 57.0 36.8 54.4 36.8 51.8 31.6 (d) General abuse 71.4 40.0 59.0 28.6 70.5 34.3 69.5 28.6 67.6 18.6 61.9 21.4 Qwen2.5-7B (SFT)Llama-3.1-8B (SFT)InternLM3-8B (SFT)GPT-5.5 (ZS)Qwen3.7-Max (ZS)DeepSeek-V4-Pro (ZS) H (harmful)NH (non-harmful) Figure 4: JREM across four context-dependent lexical-cue families. Solid and hatched bars denote harmful (H) and non-harmful (NH) cases; overlapping families are evaluated independently. Target (T) Target type (TY) Target explicitness (T) Stance (S) Label (L) Category (C) Qwen2.5-7B CoT-SFT Llama-3.1-8B CoT-SFT InternLM3-8B CoT-SFT GPT-5.5 Zero-shot Qwen3.7-Max Zero-shot DeepSeek-V4-Pro Zero-shot 23.411.615.85.93.59.3 26.814.417.67.44.910.8 24.211.214.86.43.89.2 26.910.912.14.62.410.4 32.613.915.56.13.19.8 34.814.918.99.94.810.8 (a) Field mismatch rate (%) 0 10 20 30 40 50 60 70 Mismatch rate (%) Referent (T, TY, T) Stance (S) Decision (L, C) 31.23.05.8 33.73.66.2 31.23.55.8 34.12.67.9 40.63.36.4 41.45.26.2 (b) Parse--JREM gap attribution (p) 0 10 20 30 40 Shapley contribution (p) Figure 5: Complete-record error analysis on the 1,600-item test set. (a) Field mismatch rates, with target F1 below 0.5 counted as a mismatch; (b) Shapley decomposition of the Parse–JREM gap across referent, stance, and decision fields. HER-C remains at 51.7–79.8%. CoT-SFT improves the com- plete record more consistently, raising JREM to 56.0–59.5% and lowering HER-C to 34.4–37.2%. Parse Success is nearly 100% for the CoT-SFT systems, so most remaining errors concern field content. A model may select the correct cat- egory while assigning it to the wrong target, explicitness, or author stance. Label-level scores therefore provide only a partial account of complete-record quality. Q2: How sensitive are models to lexical cues? Lexical cues alone do not determine whether a text is harmful, but strongly abusive terms can influence model judgments. We construct four challenge subsets from the test set, covering demographic identity, social status or role, body, health or disability, and general abuse. Each subset contains approxi- mately 100–200 instances with a controlled 60:40 harmful- to-non-harmful ratio. As shown in Figure 4, harmful cases achieve higher JREM in all 24 model–subset comparisons, with a median gap of 26.0 percentage points. General-abuse cues produce the largest mean gap at 38.1 points, whereas the body, health, or disability subset has the lowest overall JREM for both harmful and non-harmful cases. CoT-SFT and zero-shot API systems perform similarly on harmful cases, averaging 67.5% and 66.4% JREM, respectively. Their non- harmful JREM differs more clearly, reaching 45.8% for CoT- SFT and 35.3% for the API systems. The advantage of CoT- SFT is therefore concentrated in quotations, rebuttals, neutral mentions, and other non-harmful uses of salient cues. These results show that surface cues remain particularly difficult when their presence does not express the author’s own at- tack, while structured supervision improves the consistency of the resulting moderation records. Q3: Which fields limit complete-record prediction? Fig- ure 5 examines bottlenecks at two complementary levels. At the field level, the target has the highest mismatch rate for every system, followed by target explicitness, whereas the harmfulness label is comparatively stable. This ordering mat- ters because target localization determines the referent whose type, explicitness, stance relation, and category must then be interpreted. A boundary or normalization error can therefore coincide with several downstream mismatches even when the final decision is correct. The exact Shapley decomposition accounts for this overlap rather than counting each field in- dependently; it attributes 76.6–80.6% of the Parse Success– JREM gap to the Referent group. The field and group analyses thus converge on target localization as the main scored bot- tleneck. At the same time, the audit panel shows that reason- able readers do not always choose the same normalized target or explicitness value as the frozen reference. The measured gap consequently reflects both model difficulty and a single- answer convention. Better referent resolution, acceptance of equivalent target descriptions, and reviewer correction are therefore more appropriate remedies than treating every dis- Input:小县城公务员也就靠个编制,离了编制啥也不是。 Small-town civil servants just rely on their government tenure; without it, they are nothing. Target:小县城公务员 (small- town civil servants) Target_Type:group Target_Explicitness:explicit Stance:attack_or_endorse Label:harmful Category:occupation attack Toxic Type:hate Targeted Group:Not covered Expression:Explicitness Target:小县城公务员(small- town civil servants) Argument:are nothing without government tenure Hateful:hate Group:Others ToxiCN-styleSTATE-ToxiCNVARM-Bench Input:有人说“河南人都爱骗人”,这种地域黑真的很恶心,别再传播了。 Someone said, 'Henan people are all liars.' This kind of regional stereotyping is really disgusting. Stop spreading it. Target:河南人 (Henan people) Target_Type:group Target_Explicitness:explicit Stance:oppose_attack Label:non-harmful Category:none Target:河南人 (Henan people) Argument:deceive others Hateful:hate Group:Region Toxic Type:hate Targeted Group:regional bias Expression:Reporting STATE-ToxiCNToxiCN-style VARM-Bench Case A: Explicit status/role attack Case B: Quoted abuse with opposition Why VARM-Bench? A. More complete. Not only category; Not only target span; VARM-Bench checks the whole decision: who is targeted, how it is framed, and whether the final label is justified. Why VARM-Bench? B. More accurate. Quoted hate is not always author hate; Stance prevents false harmful decisions. Figure 6: Qualitative comparison of moderation representations. Case A shows that label/category and target-span outputs omit target properties and the linked decision basis. Case B shows that identifying quoted abuse without author stance can invert harmfulness. VARM-Bench makes all six decisions jointly inspectable. agreement as an unambiguous semantic failure. The rationale audit remains separate from reference labels and JREM. Q4: Can generated CoTs support human review? To assess whether generated CoTs are suitable for human in- spection, we randomly sample 300 shared test inputs without replacement from Qwen2.5-7B CoT-SFT and DeepSeek-V4- Pro Zero-Shot, yielding 600 CoTs. We propose three qual- ity dimensions: Groundedness checks whether the ratio- nale introduces source-unsupported claims; Adequacy re- quires sample-specific justification rather than field restate- ment; and Coherence tests consistency within the rationale and with the predicted record. Each CoT undergoes three blinded GPT-5.5 review passes, followed by human recheck- ing of failed, non-unanimous, and sampled unanimous items. CoT-SFT obtains 91.3%, 99.7%, and 99.7% on the three di- mensions, whereas Zero-Shot obtains 98.7%, 22.8%, and 100.0% (Figure 7). These results show that grounded and coherent rationales can remain too generic for inspection, while CoT-SFT more consistently provides sample-specific explanations. Generated CoTs can thus serve as inspectable review drafts, with human retaining final authority. Case Study Figure 6 shows why complete-record evaluation matters. In Case A, all representations recognize a status-based attack, but label/category output omits the affected group, while target–argument extraction omits explicitness and the link from author stance to the final decision. The VARM-Bench record makes these dependencies visible by linking the ex- plicit group target to an attacking stance, a harmful label, and the status/occupation category. Case B exposes a dif- ferent failure. The text quotes a regional stereotype only to reject it; representations without author stance can treat the Qwen2.5-7B · CoT-SFT 25 50 75 90 95 100 Adequacy Groundedness Coherence DeepSeek-V4-Pro · Zero-Shot 25 50 75 90 95 100 Adequacy Groundedness Coherence Figure 7: Generated-CoT audit profiles for Qwen2.5-7B CoT- SFT and DeepSeek-V4-Pro Zero-Shot. The three axes show Groundedness, Adequacy, and Coherence across 600 CoTs from 300 shared inputs after three blinded GPT-5.5 passes and human review. embedded claim as the author position and produce a false harmful decision. VARM-Bench retains the quoted group as the target but records opposition, non-harmfulness, and no harm category. Plausible fields can still form an inconsistent record. Together, target extraction alone cannot resolve deci- sions involving speaker commitment or discourse function. Conclusion and Future Outlook In this paper, we present VARM-Bench, a comprehensive benchmark for evaluating field-anchored rationales in Chi- nese abusive-speech moderation. The benchmark recon- structs complete moderation records from generated CoT rationales and evaluates multiple aspects of moderation deci- sions, including target identification, target properties, author stance, harmfulness labels, and categories through determin- istic metrics. Our results demonstrate that strong label-level performance can coexist with substantial errors in the under- lying moderation records, with target identification emerging as the primary source of failure. Our benchmark enables sys- tematic analysis of complete moderation decisions and model failure patterns. Future work will extend VARM-Bench to ad- ditional domains and diverse moderation policies, improve the handling of semantically equivalent target descriptions, and investigate cross-lingual transfer. Broader human evalu- ation can further assess the grounding and contextual validity of generated rationales, providing deeper insights into their reliability in real-world moderation scenarios. References Bai, Z.; Yang, L.; Yin, S.; Lu, J.; Zeng, J.; Zhu, H.; Sun, Y.; and Lin, H. 2025. STATE ToxiCN: A Benchmark for Span-level Target-Aware Toxicity Extraction in Chinese Hate Speech Detection. In Che, W.; Nabende, J.; Shutova, E.; and Pilehvar, M. T., eds., Findings of the Association for Com- putational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, volume ACL 2025, 10206–10219. Associa- tion for Computational Linguistics. Cao, Y. T.; Domingo, L.; Gilbert, S. A.; Mazurek, M. L.; Shilton, K.; and I, H. D. 2023. Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators. CoRR, abs/2311.07879. Davidson, T.; Warmsley, D.; Macy, M. W.; and Weber, I. 2017. Automated Hate Speech Detection and the Problem of Offensive Language. In Proceedings of the Eleventh In- ternational Conference on Web and Social Media, ICWSM 2017, Montréal, Québec, Canada, May 15-18, 2017, 512– 515. AAAI Press. Deng, J.; Zhou, J.; Sun, H.; Zheng, C.; Mi, F.; Meng, H.; and Huang, M. 2022. COLD: A Benchmark for Chinese Offensive Language Detection. In Goldberg, Y.; Kozareva, Z.; and Zhang, Y., eds., Proceedings of the 2022 Confer- ence on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, 11580–11599. Association for Computational Linguistics. DeYoung, J.; Jain, S.; Rajani, N. F.; Lehman, E. P.; Xiong, C.; Socher, R.; and Wallace, B. C. 2020. ERASER: A Bench- mark to Evaluate Rationalized NLP Models. In Jurafsky, D.; Chai, J.; Schluter, N.; and Tetreault, J. R., eds., Proceedings of the 58th Annual Meeting of the Association for Com- putational Linguistics, ACL 2020, Online, July 5-10, 2020, 4443–4458. Association for Computational Linguistics. ElSherief, M.; Ziems, C.; Muchlinski, D.; Anupindi, V.; Sey- bolt, J.; Choudhury, M. D.; and Yang, D. 2021. Latent Hatred: A Benchmark for Understanding Implicit Hate Speech. In Moens, M.; Huang, X.; Specia, L.; and Yih, S. W., eds., Pro- ceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, 345–363. Association for Computational Linguistics. Founta, A.; Djouvas, C.; Chatzakou, D.; Leontiadis, I.; Black- burn, J.; Stringhini, G.; Vakali, A.; Sirivianos, M.; and Kourtellis, N. 2018. Large Scale Crowdsourcing and Char- acterization of Twitter Abusive Behavior. In Proceedings of the Twelfth International Conference on Web and Social Me- dia, ICWSM 2018, Stanford, California, USA, June 25-28, 2018, 491–500. AAAI Press. Golovneva, O.; Chen, M.; Poff, S.; Corredor, M.; Zettle- moyer, L.; Fazel-Zarandi, M.; and Celikyilmaz, A. 2023. ROSCOE: A Suite of Metrics for Scoring Step-by-Step Rea- soning. In The Eleventh International Conference on Learn- ing Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net. Hartvigsen, T.; Gabriel, S.; Palangi, H.; Sap, M.; Ray, D.; and Kamar, E. 2022. ToxiGen: A Large-Scale Machine- Generated Dataset for Adversarial and Implicit Hate Speech Detection. In Muresan, S.; Nakov, P.; and Villavicencio, A., eds., Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, 3309– 3326. Association for Computational Linguistics. Jiang, A.; Yang, X.; Liu, Y.; and Zubiaga, A. 2022. SWSR: A Chinese dataset and lexicon for online sexism detection. Online Soc. Networks Media, 27: 100182. Lanham, T.; Chen, A.; Radhakrishnan, A.; Steiner, B.; Deni- son, C.; Hernandez, D.; Li, D.; Durmus, E.; Hubinger, E.; Kernion, J.; Lukosiute, K.; Nguyen, K.; Cheng, N.; Joseph, N.; Schiefer, N.; Rausch, O.; Larson, R.; McCandlish, S.; Kundu, S.; Kadavath, S.; Yang, S.; Henighan, T.; Maxwell, T.; Telleen-Lawton, T.; Hume, T.; Hatfield-Dodds, Z.; Kaplan, J.; Brauner, J.; Bowman, S. R.; and Perez, E. 2023. Mea- suring Faithfulness in Chain-of-Thought Reasoning. CoRR, abs/2307.13702. Lightman, H.; Kosaraju, V.; Burda, Y.; Edwards, H.; Baker, B.; Lee, T.; Leike, J.; Schulman, J.; Sutskever, I.; and Cobbe, K. 2024. Let’s Verify Step by Step. In The Twelfth Interna- tional Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. Liu, K.; Cheng, S.; Tian, B.; Liang, X.; Yin, Y.; Han, M.; Zhang, N.; Hooi, B.; Chen, X.; and Deng, S. 2025. ChineseHarm-Bench: A Chinese Harmful Content Detec- tion Benchmark. CoRR, abs/2506.10960. Lu, J.; Xu, B.; Zhang, X.; Min, C.; Yang, L.; and Lin, H. 2023. Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and Bench- marks. In Rogers, A.; Boyd-Graber, J. L.; and Okazaki, N., eds., Proceedings of the 61st Annual Meeting of the Associa- tion for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, 16235–16250. Association for Computational Linguistics. Mathew, B.; Saha, P.; Yimam, S. M.; Biemann, C.; Goyal, P.; and Mukherjee, A. 2021. HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection. In Thirty- Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Arti- ficial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, volume 35, 14867–14875. AAAI Press. Mazeika, M.; Phan, L.; Yin, X.; Zou, A.; Wang, Z.; Mu, N.; Sakhaee, E.; Li, N.; Basart, S.; Li, B.; Forsyth, D. A.; and Hendrycks, D. 2024. HarmBench: A Standardized Evalu- ation Framework for Automated Red Teaming and Robust Refusal. In Salakhutdinov, R.; Kolter, Z.; Heller, K. A.; Weller, A.; Oliver, N.; Scarlett, J.; and Berkenkamp, F., eds., Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, volume 235, 35181–35224. PMLR / OpenReview.net. Pavlopoulos, J.; Sorensen, J.; Laugier, L.; and Androutsopou- los, I. 2021. SemEval-2021 Task 5: Toxic Spans Detec- tion. In Palmer, A.; Schneider, N.; Schluter, N.; Emerson, G.; Herbelot, A.; and Zhu, X., eds., Proceedings of the 15th International Workshop on Semantic Evaluation, Se- mEval@ACL/IJCNLP 2021, Virtual Event / Bangkok, Thai- land, August 5-6, 2021, 59–69. Association for Computa- tional Linguistics. Prasad, A.; Saha, S.; Zhou, X.; and Bansal, M. 2023. ReCE- val: Evaluating Reasoning Chains via Correctness and In- formativeness. In Bouamor, H.; Pino, J.; and Bali, K., eds., Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, 10066–10086. Association for Com- putational Linguistics. Röttger, P.; Vidgen, B.; Nguyen, D.; Waseem, Z.; Margetts, H. Z.; and Pierrehumbert, J. B. 2021. HateCheck: Func- tional Tests for Hate Speech Detection Models. In Zong, C.; Xia, F.; Li, W.; and Navigli, R., eds., Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, 41–58. Association for Computational Linguistics. Sap, M.; Gabriel, S.; Qin, L.; Jurafsky, D.; Smith, N. A.; and Choi, Y. 2020. Social Bias Frames: Reasoning about Social and Power Implications of Language. In Jurafsky, D.; Chai, J.; Schluter, N.; and Tetreault, J. R., eds., Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, 5477–5490. Association for Computational Linguistics. Schöpke-Gonzalez, A. M.; Wu, S.; Kumar, S.; Resnick, P. J.; and Hemphill, L. 2023. How We Define Harm Impacts Data Annotations: Explaining How Annotators Distinguish Hate- ful, Offensive, and Toxic Comments. CoRR, abs/2309.15827. Turpin, M.; Michael, J.; Perez, E.; and Bowman, S. R. 2023. Language Models Don’t Always Say What They Think: Un- faithful Explanations in Chain-of-Thought Prompting. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; and Levine, S., eds., Advances in Neural Information Process- ing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023. Wang, B.; Chen, W.; Pei, H.; Xie, C.; Kang, M.; Zhang, C.; Xu, C.; Xiong, Z.; Dutta, R.; Schaeffer, R.; Truong, S. T.; Arora, S.; Mazeika, M.; Hendrycks, D.; Lin, Z.; Cheng, Y.; Koyejo, S.; Song, D.; and Li, B. 2023a. DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; and Levine, S., eds., Advances in Neural Informa- tion Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023. Wang, X.; Wei, J.; Schuurmans, D.; Le, Q. V.; Chi, E. H.; Narang, S.; Chowdhery, A.; and Zhou, D. 2023b. Self- Consistency Improves Chain of Thought Reasoning in Lan- guage Models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net. Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E. H.; Le, Q. V.; and Zhou, D. 2022. Chain- of-Thought Prompting Elicits Reasoning in Large Language Models. In Koyejo, S.; Mohamed, S.; Agarwal, A.; Belgrave, D.; Cho, K.; and Oh, A., eds., Advances in Neural Informa- tion Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, vol- ume 35, 24824–24837. Yu, X.; Blanco, E.; and Hong, L. 2022. Hate Speech and Counter Speech Detection: Conversational Context Does Matter. In Carpuat, M.; de Marneffe, M.; and Ruíz, I. V. M., eds., Proceedings of the 2022 Conference of the North Amer- ican Chapter of the Association for Computational Linguis- tics: Human Language Technologies, NAACL 2022, Seattle, WA, United States, July 10-15, 2022, 5918–5930. Associa- tion for Computational Linguistics. Zampieri, M.; Morgan, S.; North, K.; Ranasinghe, T.; Simm- mons, A.; Khandelwal, P.; Rosenthal, S.; and Nakov, P. 2023. Target-Based Offensive Language Identification. In Rogers, A.; Boyd-Graber, J. L.; and Okazaki, N., eds., Proceedings of the 61st Annual Meeting of the Association for Compu- tational Linguistics (Volume 2: Short Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, 762–770. Association for Computational Linguistics. Zhang, Z.; Lei, L.; Wu, L.; Sun, R.; Huang, Y.; Long, C.; Liu, X.; Lei, X.; Tang, J.; and Huang, M. 2024. SafetyBench: Evaluating the Safety of Large Language Models. In Ku, L.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, 15537–15553. Association for Computational Linguistics. Zheng, J.; Liu, X.; Haque, M.; Qian, X.; Yang, G.; and Yang, W. 2024. HateModerate: Testing Hate Speech Detectors against Content Moderation Policies. In Duh, K.; Gómez- Adorno, H.; and Bethard, S., eds., Findings of the Association for Computational Linguistics: NAACL 2024, Mexico City, Mexico, June 16-21, 2024, volume NAACL 2024, 2691– 2710. Association for Computational Linguistics. Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E. P.; Zhang, H.; Gonzalez, J. E.; and Stoica, I. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv:2306.05685. Zhou, J.; Deng, J.; Mi, F.; Li, Y.; Wang, Y.; Huang, M.; Jiang, X.; Liu, Q.; and Meng, H. 2022. Towards Identifying Social Bias in Dialog Systems: Frame, Datasets, and Benchmarks. arXiv:2202.08011.