Paper deep dive
How Should AI Safety Benchmarks Benchmark Safety?
Cheng Yu, Severin Engelmann, Ruoxuan Cao, Dalia Ali, Orestis Papakyriakopoulos
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/11/2026, 1:07:03 AM
Summary
This paper reviews 210 AI safety benchmarks, identifying significant technical, epistemic, and sociotechnical shortcomings. It proposes a framework for improving safety benchmarking by applying risk management principles from engineering, such as the Rumsfeld matrix for construct coverage, probabilistic risk assessment for quantification, and measurement theory for validity.
Entities (5)
Relation Signals (4)
AI Safety Benchmarks → haslimitation → Construct Coverage Imbalance
confidence 95% · 81% of surveyed benchmarks evaluate only predefined known risks
AI Safety Benchmarks → lacksrigor → Risk Quantification
confidence 95% · 79% of benchmarks reduce safety to binary pass/fail rates
Measurement Theory → ensures → Measurement Validity
confidence 90% · we apply measurement theory to ensure epistemologically sound construct definitions
Rumsfeld Matrix → improves → Construct Coverage
confidence 90% · For construct coverage, we apply the Rumsfeld matrix of known/unknown risks to systematically map blind spots
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:AI safety benchmarks are pivotal for safety in advanced AI systems; however, they have significant technical, epistemic, and sociotechnical shortcomings. We present a review of 210 safety benchmarks that maps out common challenges in safety benchmarking, documenting failures and limitations by drawing from engineering sciences and long-established theories of risk and safety. We argue that adhering to established risk management principles, mapping the space of what can(not) be measured, developing robust probabilistic metrics, and efficiently deploying measurement theory to connect benchmarking objectives with the world can significantly improve the validity and usefulness of AI safety benchmarks. The review provides a roadmap on how to improve AI safety benchmarking, and we illustrate the effectiveness of these recommendations through quantitative and qualitative evaluation. We also introduce a checklist that can help researchers and practitioners develop robust and epistemologically sound safety benchmarks. This study advances the science of benchmarking and helps practitioners deploy AI systems more responsibly.
Tags
Links
Trouble viewing inline? Open PDF directly →
Full Text
187,103 characters extracted from source content.
Expand or collapse full text
How Should AI Safety Benchmarks Benchmark Safety? Cheng Yu 1 Severin Engelmann 2 Ruoxuan Cao 1 Dalia Ali 1 Orestis Papakyriakopoulos 1 Abstract AI safety benchmarks are pivotal for safety in advanced AI systems; however, they have sig- nificant technical, epistemic, and sociotechnical shortcomings. We present a review of 210 safety benchmarks that maps out common challenges in safety benchmarking, documenting failures and limitations by drawing from engineering sciences and long-established theories of risk and safety. We argue that adhering to established risk man- agement principles, mapping the space of what can(not) be measured, developing robust proba- bilistic metrics, and efficiently deploying measure- ment theory to connect benchmarking objectives with the world can significantly improve the valid- ity and usefulness of AI safety benchmarks. The review provides a roadmap on how to improve AI safety benchmarking, and we illustrate the effectiveness of these recommendations through quantitative and qualitative evaluation. We also introduce a checklist 1 that can help researchers and practitioners develop robust and epistemo- logically sound safety benchmarks. This study advances the science of benchmarking and helps practitioners deploy AI systems more responsibly. 1. Introduction The rapid advances in artificial intelligence (AI) are cre- ating systems with ever-increasing capabilities and access to diverse environments. While these developments hold huge potential for societal benefit, they also introduce risks to safety, ranging from malicious use and manipulation to malfunctions and systemic issues (Amodei et al., 2016; Brundage et al., 2018; Weidinger et al., 2022). Harms re- lated to AI have been documented broadly–in algorithmic bias and discrimination in computer vision (Buolamwini 1 Societal Computing, Technical University of Munich, Munich, Germany 2 Department of Information Science, Cornell Univer- sity, NY, US. Correspondence to: Cheng Yu <cheng.yu@tum.de>, Orestis Papakyriakopoulos <orestis.p@tum.de>. Preprint. February 10, 2026. 1 https://anonymous.4open.science/r/ ai-safety-benchmark/ and Gebru, 2018), toxicity and harmful content generation by language models (Gehman et al., 2020), and privacy leak- age via training data extraction (Carlini et al., 2021). Risks have been identified in malicious uses across digital, physi- cal, and political domains (Brundage et al., 2018), dual-use biological and chemical design (Urbina et al., 2022), and systemic reliability under distribution shift (Koh et al., 2021). The most common solution to mitigate risks has been the development and use of AI safety benchmarks (Liang et al., 2022; Vidgen et al., 2024; Center for AI Safety, 2025). This solution is an extension of the traditional benchmarking culture in computer science, where standardized tests are designed and conducted to evaluate and compare the perfor- mance of computer systems, components, and algorithms (Russakovsky et al., 2015; Standard Performance Evalua- tion Corporation, 2017; Mattson et al., 2020). Benchmarks are useful and can reveal vulnerabilities of AI systems, and have clearly contributed to the development of the field of AI both scientifically and in its application (Ribeiro et al., 2020; Gehman et al., 2020; Koh et al., 2021; Zou et al., 2023). However, the discipline of benchmarking in AI safety— understood here as “endeavor dedicated to preventing or mitigating harms from AI systems” (Harding and Kirk- Giannini, 2025a)—overlooks established safety-related the- ories, frameworks, practices, and knowledge developed over past decades for modeling, measuring, and mitigating risk (IEC, 2010; Leveson, 2011a; Hollnagel, 2014; ISO, 2018; 2023; NIST, 2023). Given this, we answer: What are the limitations of AI safety benchmarks? How can we leverage existing theories, frameworks, and practices of safety and safety engineering to improve AI safety benchmarks? Drawing on a review of 210 AI safety benchmarks, we identify three core limitations. First, construct coverage is imbalanced: 81% of surveyed benchmarks evaluate only predefined known risks (e.g., toxicity or jailbreaks via fixed prompts), leaving emergent behaviors and unforeseen fail- ures unexamined. Second, risk quantification lacks prob- abilistic rigor: 79% of benchmarks reduce safety to binary pass/fail rates, treating empirical frequencies as calibrated probabilities while ignoring severity. Third, measurement validity erodes through proxy chains: metrics like refusal rates are conflated with real-world outcomes, yet halving a toxicity score does not necessarily halve actual harm. To answer how we can leverage existing theories, frameworks, 1 arXiv:2601.23112v2 [cs.CY] 8 Feb 2026 How Should AI Safety Benchmarks Benchmark Safety? and practices of safety and safety engineering to improve AI safety benchmarks, we draw on three established bodies of knowledge to propose ten recommendations (R1–R10). For construct coverage, we apply the Rumsfeld matrix of known/unknown risks to systematically map blind spots and prioritize discovery of novel failure modes (R1–R3). For risk quantification, we adopt probabilistic risk assess- ment, replacing binary frequencies with calibrated proba- bilities and operationalizing risk as severity×likelihood (R4–R6). For measurement validity, we apply measure- ment theory to ensure epistemologically sound construct definitions, traceable calibration, and deployment-grounded proxies (R7–R10). We provide quantitative and qualitative illustrations for translating benchmark scores to deployment risk (App. C) and a benchmark design checklist (App. D). Together, these operationalize safety benchmarking as a normative process connecting abstract values to real-world outcomes. Complete coding results are reported in App. E. 2. The Uniqueness of Safety Benchmarking To understand the limitations of existing benchmarks and how to improve them, it is necessary to identify what makes safety benchmarking distinct. Traditionally, benchmarks are fixed test sets created using holdout methods and reused to ensure comparable evaluation (Hardt and Recht, 2021). So- cially, a benchmark is a community framework combining datasets with a metric aligned to a technical task. This met- ric aggregates performance into a single score, where high- scoring models are considered state-of-the-art (Raji et al., 2021). This often involves leaderboards (Orr and Kang, 2024) to recognize technical achievements and encourage competition. These descriptions yield two observations: benchmarking has historically been linked with maximizing capability, and it focuses on technical objectives reflecting the latest technological advancement. AI safety benchmarks differ fundamentally from traditional evaluations by focusing on risk mitigation rather than task proficiency (Center for AI Safety, 2025). This shift in- volves two critical dimensions: normative assessment and sociotechnical context. Safety benchmarks are normative rather than descriptive. Traditional benchmarks measure how well a model performs, while safety benchmarks as- sess potential to cause harm (Buolamwini and Gebru, 2018). Descriptively, a model like GPT-5 outperforms GPT-2 in co- herence and knowledge. Normatively, GPT-5 may be judged worse because its capabilities enable harmful outputs such as weapon design that GPT-2 simply cannot produce. Safety benchmarks are also sociotechnical rather than purely tech- nical (Dobbe, 2022). GPT-5 excels technically by accom- plishing more tasks, but what counts as “better” depends on human usage. The same capability that enables chem- istry tutoring also facilitates harm. Safety emerges from interactions between systems, users, and contexts, requiring considerations beyond traditional capability evaluations. 3. Risk & Safety in Engineering vs Benchmarking To construct AI safety benchmarks that truly reflect the field’s normative and sociotechnical nature, we can look to the operationalization of "risk" and "safety" in long estab- lished scientific fields. In safety engineering and risk man- agement, risk measurement functions as a two-step process. First, risk measurement bridges the gap between abstract social values and physical reality. It operationalizes risk not merely as technical failure, but as a function of the mag- nitude of consequences, mediated by hazards and system vulnerabilities (IPCC, 2012; ISO, 2018; NIST, 2023). This step translates normative concepts, such as what constitutes “harm” or “vulnerability”, into concrete observable phenom- ena. Second, to maintain this link despite the complexity of real world and the uncertain manifestation of harm, risk mea- surement employs probability theory. This is concretized in functional safety, where “acceptable” risk thresholds are instantiated as target probabilities of dangerous failure (e.g., Safety Integrity Levels) (IEC, 2010; Leveson, 2011b). By quantifying and qualifying the likelihood and consequence of these events, safety engineering reduces real-world uncer- tainty to a manageable metric, ensuring that systems, from commonplace applications to high-stakes systems, operate within socially accepted bounds (Dezfuli et al., 2011; ISO 14971:2019; European Union Agency for Railways, 2022). Translating this engineering-based conception to AI safety implies that robust safety benchmarks must successfully ex- ecute both steps: connecting normative values to real-world indicators, and handling uncertainty through probability. However, it is not clear to what extent AI safety efforts achieve the above objectives. There is indeed a vast amount of frameworks that attempt to map the theoretical landscape of safety. For example, Weidinger et al. (2022) and the International AI Safety Report (DSIT, 2025) provide cata- logs of normative harms ranging from bias to systemic risks, while HELM (Liang et al., 2022) emphasizes broad scenario coverage. Guided by these, benchmarks like TruthfulQA (Lin et al., 2022), MACHIAVELLI (Pan et al., 2023), and HarmBench (Mazeika et al., 2024) attempt to measure these values. Nonetheless, recent analyses (Zhao et al., 2024a; Bean et al., 2025) on construct validity highlight a criti- cal gap: benchmarks often fail to establish a clear connec- tion between what they claim to measure (e.g., "diversity", "safety") and what their metrics actually capture. Further- more, current benchmarks rely on distinct metrics—such as refusal rates, keyword matching, or attack success—which differ significantly from the actual manifestation of harm (Jacobs and Wallach, 2021). They also treat safety as a 2 How Should AI Safety Benchmarks Benchmark Safety? static checklist of “known knowns” rather than a probabilis- tic function of uncertainty (Leveson, 2011b; Amodei et al., 2016). By focusing on fixed metrics, they neglect the "like- lihood" and “severity” calculus central to risk management (NIST, 2023). This reliance on deterministic metrics not only ignores the engineering definition of risk but invites metric gaming and target fixation, exemplifying Goodhart’s and Campbell’s laws (Goodhart, 1975; Campbell, 1979). Thus, to move beyond theoretical critique, quantify these methodological gaps and develop actionable recommenda- tions, we perform a comprehensive survey of the existing AI safety benchmarking landscape. 4. Method We conduct an extensive scoping review of literature on AI safety benchmarks. To classify benchmarks as AI safety related, we adopt the definition of AI safety as the “en- deavor dedicated to preventing or mitigating harms from AI systems” (Harding and Kirk-Giannini, 2025a), detailed in App. A.1. Our goal is to map the terrain of AI safety bench- marks comprehensively, applying a broad lens of analysis across three key dimensions, elaborated in Sections 5–7: 1) the types of risks AI safety benchmarks are designed to detect; 2) how do they quantify risks and harms; 3) how do they ensure what they measure links to the world. Fig. 1 summarizes key concerns and recommendations; detailed framework discussion underlying these evaluation dimen- sions is provided in App. B. Throughout our analysis, we examine the engagement with the sociotechnical nature of safety, recognizing that safety emerges from interactions between AI systems, users, and societal contexts. 5. Dealing with Safety Construct Coverage The Rumsfeld matrix shown in Fig. 2 offers a useful lens for examining which risks AI safety benchmarks evaluate and which they systematically neglect. Following previous work on AI safety engineering (Wisakanto et al., 2025), we adapt the matrix along two epistemic dimensions: Aware- ness (whether we are conscious of a risk) and Understanding (whether we possess empirical knowledge or verified failure modes). This yields four quadrants: known knowns (empiri- cally verified risks we actively monitor), known unknowns (anticipated emergent behaviors we do not yet fully under- stand), unknown knowns (theoretical risks or documented phenomena not currently identified in practice), and un- known unknowns (entirely unforeseen behaviors for which no prior data exists). Detailed definitions and examples dis- tinguishing these categories are provided in App. B.1. As shown in left panel of Fig. 1, mapping 210 benchmarks to this uncertainty framework reveals a pronounced imbalance. Known knowns dominate (N=170). Most benchmarks evaluate predefined risk types with predetermined trig- gers, such as bias measurement through demographic tem- plates and jailbreak evaluations with fixed adversarial exam- ples. Known unknowns receive limited attention (N=33). Benchmarks like GPTFuzz (Yu et al., 2024a) and WildTeam- ing (Jiang et al., 2024a) search for novel instantiations of understood risks through fuzzing and red-teaming, yet tool- use vulnerabilities and multi-step reasoning failures remain largely unexplored. Unknown knowns are largely ignored (N=5). Well-documented ML phenomena, including distri- bution shift (Filos et al., 2020; Zhang et al., 2025a), out-of- distribution detection (Goodier and Campbell, 2023), and differential harms to vulnerable populations (Berman and Albright, 2017), rarely transfer from robustness and ethics research into safety evaluation frameworks. Unknown un- knowns remain nearly absent (N=2). Rare exceptions such as Perez et al. (2023) and LLMArena (Chen et al., 2024a) demonstrate that unanticipated risks, including inverse scal- ing, emergent goal-seeking, and multi-agent herding, are discoverable through appropriate methodology. Nonethe- less, investment in developing such approaches remains limited across the broader research community. This distribution creates structural blind spots: systems opti- mized for anticipated risks often remain vulnerable to unan- ticipated ones (Taleb, 2007). Emphasis on known knowns risks unwarranted confidence. Recommendations below aim to narrow these coverage gaps. R1. Documenting Known Blind Spots. Effective safety benchmarks benefit from a limitations section specify- ing which risk types are covered versus excluded, along with assumptions about deployment context. Only 34% (N=72) of surveyed benchmarks explicitly specify the risks they uncover, such as data contamination (Gupta et al., 2024a), the complex and socially constructed nature of tasks (Laszkiewicz et al., 2024), or other vulnerabilities and hypothetical real-world harms (Souly et al., 2024). Iden- tifying blind spots upfront rather than discovering them post-deployment helps prevent over-interpretation of bench- mark scores as comprehensive safety assessments. For in- stance, a jailbreak benchmark (Yu et al., 2024b) that explic- itly acknowledges its exclusion of multi-turn manipulation or context-dependent harm amplification provides clearer guidance for practitioners assessing deployment readiness. R2. Expanding Known Boundaries. Current bench- marks predominantly rely on predefined prompts that target well-understood abstractions and are optimized for ease of measurement. Discovering risks beyond this bounded de- sign space calls for sustained investment in progressively open-ended evaluation methods. Algorithmic approaches such as automated fuzzing (Zhang et al., 2023) and self- evolving reframing operations (Wang et al., 2025a) system- atically stress-test guardrails by transforming seed prompts 3 How Should AI Safety Benchmarks Benchmark Safety? SAFETY CONSTRUCTS Types of risks benchmarks are designed to detect R1. Documenting Known Blind Spots R2. Expanding Known Boundaries R3. Reframing Known ML Phenomena as Safety Concerns RISK QUANTIFICATION How should risks be numerically represented? R4. Calibrating Benchmark Frequencies to Exposure R5. Grounding Severity in Principled Frameworks R6. Accounting for Uncertainty Quantification C5. Empirical frequencies misused as probabilities C6. Severity scales lack principled grounding MEASUREMENT THEORY Ensuring measurements link to real-world safety R7. Standardizing Safety Constructs with Transparency R8. Locking and Versioning for Reproducibility R9. Anchoring Proxies in Deployment Contexts R10. Iterative Refinement via Community Input C7. Unstandardized metrics prevent comparison C8. Lack of accuracy and precision C9. Construct validity erodes through proxy chains. C1. Known knowns dominate (81%) C4. Unknown unknowns nearly absent (1%) C2.Known unknowns with limited attention (16%) C3. Unknown knowns issues are ignored (2%) Figure 1. Framework for improving AI safety benchmarking. Based on an analysis of 210 benchmarks, the figure summarizes key concerns (C1–C9) and recommendations (R1–R10) across three dimensions: expanding coverage of safety constructs beyond known knowns, adopting principled risk quantification with probabilistic rigor, and aligning measurements with real-world safety outcomes. Known Knowns What we know andare aware that we know e.g.,“We know our system handles 10,000 users.” Known Unknowns What we know we don’t know e.g., “We don’t know how users will behave during peak traffic.” Unknown Knowns Things we know but don’t realize we know e.g., A team member has critical experience but is never asked Unknown Unknowns Things we don’t know and haven’t thought to ask e.g., A failure mode no one imagined because the question was never raised Things we DON’T Know Things we Know Questions we Ask Questions we DON’T Ask Level of Understanding Level of Awareness + - - + Figure 2. Rumsfeld matrix mapping awareness and understanding. into increasingly complex syntactic or semantic variants. Beyond algorithmic evolution, uncovering unanticipated behaviors benefits from exploratory and participatory ap- proaches. These include scalable evaluations where Perez et al. (2023) uses LM-generated evaluation to discover novel failure modes, as well as multi-agent stress testing to reveal emergent risks such as herding behavior or bias amplifica- tion that arise only through interaction (Chen et al., 2024a; Yuan et al., 2024). Institutionalizing these discoveries in- volves establishing community contribution mechanisms, including validated red-teaming submission portals with versioned integration and contributor credit, alongside par- ticipatory design (Google DeepMind, 2025) that engages external stakeholders to surface otherwise invisible harms. R3. Reframing Known ML Phenomena as Safety Con- cerns. Many well-understood machine learning problems, typically discussed only in specialized research, warrant recognition as safety concerns. This reframing broadens the scope of responsibility by aligning technical evalua- tion with operational, societal, and ethical consequences, thereby motivating stronger standards for evaluation and oversight. Distribution shift offers a clear case: when a model encounters data that differs from its training distri- bution, performance degrades. CARNOVEL (Filos et al., 2020) frames this generalization failure as safety-critical rather than a mere performance limitation. Annotation bias presents a subtler challenge. Systematic distortions can emerge from annotator selection, disagreement patterns, and demographic skew, quietly privileging certain perspectives over others. Data contamination poses yet another risk. As models train on increasingly comprehensive internet data, they may have already seen nominally held-out test exam- ples during training. Addressing these concerns involves implementing contamination detection, track temporal va- lidity, and design evaluation sets that resist leakage. 6. Benchmarking Safety via Risk Attributes Safety engineering characterizes risk through two core at- tributes: explicit probabilities of violation and severities of consequence (NIST, 2023). This section examines existing AI safety benchmarks through this lens, focusing on how they define or approximate violation likelihood and how they represent outcome severity. Across the benchmarks surveyed, neither dimension is instantiated in a way that yields calibrated or decision-relevant measures of risk. Empirical frequencies misused as calibrated probabili- ties. 79% (N=166) of surveyed benchmarks rely on binary outcome proportions as their primary or sole evaluation metric, reducing safety assessment to pass–fail rates. This pattern recurs across safety domains, including bias evalua- tion (biased/unbiased) (Nangia et al., 2020; Dhamala et al., 2021; Parrish et al., 2021), adversarial robustness (attack 4 How Should AI Safety Benchmarks Benchmark Safety? success/failure) (Wang et al., 2021; Luo et al., 2024; Yu et al., 2024a), and general harm assessment (harmful/harm- less) (Ghosh et al., 2025; Mazeika et al., 2024; Li et al., 2024a). While facilitating ease of operationalization, this methodological uniformity tends to obscure variation in severity and contextual dependence. Meanwhile, bench- marks risk a conceptual misalignment by presenting these empirical frequencies as “probabilities” of unsafe behav- ior (Zhang et al., 2023; Hall et al., 2023). Probabilistic risk assessment (PRA) in safety engineering, by contrast, treats probability as an estimate that incorporates uncertainty, en- vironmental variability, and dependencies among failure modes. Current AI safety benchmarks instead typically report point estimates without confidence intervals, uncer- tainty modeling, or robustness to distributional shift. The resulting quantities remain disconnected from the causal structure and epistemic rigor that meaningful risk character- ization requires. Severity scales lack principled grounding. When bench- marks move beyond binary labels, they frequently adopt ordinal severity scales of harm (e.g., 1–5 or A–F) without clear justification of their cardinal interpretation or norma- tive basis (Jiang et al., 2024b; Dineen et al., 2025). Whether adjacent levels correspond to comparable increments of harm often remains unclear, limiting interpretability and undermining cross-benchmark comparisons. Of the 210 benchmarks surveyed, only 36 distinguish between lev- els of harm severity; among these, just 14 provide prin- cipled justification for these distinctions, drawing on prior research (Bianchi et al., 2023), industry standards (Han et al., 2024), or AI usage policies (Shen et al., 2024a). The re- mainder rely on ad hoc author judgment (Souly et al., 2024), LLM-generated labels (Dineen et al., 2025), or provide no stated rationale (Huang and Xiong, 2024). These gaps suggest several directions for improvement. R4. Calibrating Benchmark Frequencies to Exposure Estimates. Although the AI safety benchmarking commu- nity widely notes that no evaluation guarantees “absolute” safety (Ghosh et al., 2025), this caution is not always re- flected in quantitative reporting. Current language conflates what benchmarks measure, namely resistance to specific prompts, with safety in general. We therefore recommend using empirically grounded terms, e.g., empirical rate, ob- served frequency, or sample proportion, rather than proba- bility , which can invite overgeneralization. Relatedly, met- rics such as perplexity and token likelihood are sometimes used to motivate probabilistic readings (Zhao et al., 2023); however, they primarily quantify a model’s relative fit to particular sequences and do not directly provide calibrated generation probabilities or deployment-level estimates of adverse-event risk. In addition, calibrating benchmark rates using in-the-wild prevalence estimates may support more risk-relevant interpretation. As shown in App. C.1, com- bining benchmark failure rates with real-world prevalence indicates that systems with similar benchmark scores can differ by an order-of-magnitude in implied deployment risk. R5. Grounding Severity in Principled Frameworks. Graded severity scales benefit from explicit justification rather than ad hoc author judgments or LLM-generated ratings. Justifications may draw on prior empirical re- search (Bianchi et al., 2023), domain-specific normative frameworks such as medical ethics (Pal et al., 2023), or regulatory classifications (Zeng et al., 2024). Clarifying whether severity levels represent equal intervals, power-law relationships, or catastrophic thresholds enhances meaning- ful comparison. This parallels practices in safety-critical fields such as the Common Vulnerability Scoring System in cybersecurity and the Abbreviated Injury Scale in trauma medicine. We illustrate this in App. C.2 with an order- of-magnitude calculation, following estimation tradition of approximate reasoning with limited data (Fermi, 1945). By translating benchmark failure rates into expected mon- etary losses through prevalence propagation and empirical severity distributions, we show that annual liability for a medium-sized platform is on the order of10 4 $under typical severity assumptions, revealing deployment risks invisible to raw benchmark scores. R6. Accounting for Uncertainty Quantification. 94% (N=198) benchmarks acknowledge uncertainty, typically via disclaimers, including evaluator uncertainty, model in- stability , and data sampling uncertainty. Practical mitiga- tion efforts largely focus on the reliability of evaluators or human annotations, operationalized via voting schemes and inter-annotator agreement metrics. Some works report in- sample uncertainty measures, such as worst-case bounds derived from concentration inequalities (Zhang et al., 2018) or 95% confidence intervals (Souly et al., 2024). For com- putationally expensive sources of uncertainty, limited work explores solutions such as multi-run evaluations (Kim et al., 2022) or testing prompt variations (Wang et al., 2025a). Beyond uncertainty within the test distribution, conceptual and normative uncertainty remains in how benchmark per- formance maps to deployment risk, which is acknowledged only via scope disclaimers. Engineering disciplines employ multiplicative safety factors as a form of conservative rea- soning (Dourson and Stara, 1983). Extrapolating benchmark results to deployment-level risk may incorporate analogous safety margins to account for coverage gaps, distributional shift, and model instability. These margins could be in- formed by conservative bounds based on domain-specific risk tolerance or expert elicitation regarding plausible failure amplification in deployment. Under this view, benchmark failure rates serve as lower bounds on deployment risk, with safety margins communicating residual uncertainty. 5 How Should AI Safety Benchmarks Benchmark Safety? 7. Aligning Safety with Measurement Theory Even when benchmarks are accepted as necessary proxies for real-world safety, their measurement practices often violate core principles from measurement science. Across the AI safety benchmarks reviewed, three critical limitations emerge that are especially acute for safety evaluation. Unstandardized metrics prevent meaningful safety claims. In mature measurement sciences, reliable evalua- tion depends on standardization grounded in proportionality, invariance, and traceable calibration (Tal, 2020). For ex- ample, temperature measurements are comparable because their scales are anchored to physical reference points such as freezing point of water. AI safety metrics lack such empirical grounding: only 38% (N=79) explicitly ground definitions or proxies in established framework, external regulations or societal standards. Scoring schemes rarely specify what real-world quantity they approximate, whether score differences correspond to proportional changes in risk, or how metrics behave across deployment settings. For ex- ample, halving a toxicity score (e.g., 0.48→0.23) does not necessarily halve user exposure to harm, as the scale is typi- cally unvalidated and its relationship to real-world outcomes remains unknown (Gehman et al., 2020). Few benchmarks attempt traceable calibration; MEDFAIR (Zong et al., 2023) is a notable exception, linking fairness metrics to established clinical performance measures. The absence of standard- ization limits comparability across studies and complicates deployment decisions, as practitioners lack clear guidance on what benchmark scores imply about real-world safety. Lack of Accuracy and Precision. Accuracy refers to close- ness to the true value, while precision concerns the sta- bility of repeated measurements (Tal, 2020). Current AI safety benchmarks struggle to achieve either property. Many benchmarks report variance (e.g., mean ± sd toxicity scores), but these metrics reflect only internal instability. The “truth value” they approximate is typically automated labeler or LLM judge. Without calibration against field outcomes, such numbers fail to track real-world harm. Precision is similarly limited: scores frequently vary with random seeds, prompt phrasing, or evaluator versions. Yet 89% (N=186) of benchmarks evaluate on pre-defined fixed data-without doc- umented sources of stochasticity, apparent improvements are difficult to distinguish from measurement noise. Construct validity erodes through proxy chains. Con- struct validity concerns whether a score serves as a defensi- ble proxy for the real phenomenon of interest. Recent work applying measurement theory to AI evaluation highlights pervasive validity failures: unclear constructs, mismatched measurements, and limited justification for why metrics cap- ture target constructs (Bean et al., 2025; Salaudeen et al., 2025; Wallach et al., 2025). In AI safety evaluation, these issues are compounded by a proxy-of-a-proxy structure: ab- stract safety constructs are first operationalized through benchmark scenarios or prompts, and then further reduced to model outputs and numerical scores. The conceptual complexity of safety constructs poses unique validity challenges. 68% (N=143) of surveyed benchmarks rely on isolated, single-turn model interactions, diverging from how AI systems function in real-world safety-critical settings. Unlike capability constructs (e.g., Olympiad math or GitHub coding capability (Petrov et al., 2025; Jimenez et al., 2023)) that are contested but bounded, safety-relevant uses of AI are highly contextual, interactive, and embed- ded in social institutions. Sociotechnical systems research has documented several pitfalls of abstraction such as the formalism trap (Selbst et al., 2019; Dobbe, 2022). In prac- tice, many of the harms addressed within the AI safety discourse exist only in relation to competing values and interests. However, these value conflicts surface only when looking at the specific contexts in which they are placed. Applying an open conceptualization of harms on the one hand, while instantiating this narrow perspective of oper- ationalizing safety in testing on the other, inevitably gen- erates gaps between what the benchmark purports to the test and what conceptualization of a contested concept it actually measures. The safety construct remains under- specified, and consequently its formalization. For instance, Mazeika et al. (2024) reports a single attack success rate aggregated across diverse semantic categories, obscuring qualitative differences in risk. Detailed discussion of va- lidity challenges, including contextual value conflicts and benchmark-deployment gaps, appears in App. B.3.1. R7. Standardizing Safety Constructs with Transparency. “Safety” is not a unitary concept, and meaningful measure- ment benefits from grounding in core principles of measure- ment science. Benchmarks can improve clarity by spec- ifying the harm constructs they target (e.g., toxicity, bias, manipulation) and providing operational definitions for each. Transparency is further improved by stating whose values inform judgments of harm, such as expert assessments, pol- icy frameworks, or affected communities, and by acknowl- edging contested normative choices. Finally, articulating the relationship between measured proxies and real-world safety concerns supports more informed interpretation. As illustrated in App. C.1, transforming model-centric scores into deployment-grounded exposure estimates offers one example of traceable calibration that connects benchmark outputs to core measurement-theoretic principles. R8. Locking and Versioning for Reproducibility. Repro- ducibility ensures a rigorous benchmark design. Model ac- cess specifications benefit from going beyond coarse labels (e.g., “GPT-4”) by recording API endpoints, access dates, weight checksums, quantization methods, and inference parameters. Fixing and reporting sources of stochasticity, 6 How Should AI Safety Benchmarks Benchmark Safety? including random seeds for data sampling, model inference, and evaluation procedures, can help ensure consistent results across independent evaluations. Evaluation context simi- larly warrants verbatim versioning: system prompts, LLM judge versions, constitutional principles, and scoring rubrics all merit exact recording, as even minor prompt changes may substantially affect safety judgments. R9. Anchoring Proxies in Deployment Contexts. As Ris- mani et al. (2025) many AI ethics measures focus narrowly on model outputs while neglecting data quality, user expe- rience, and systemic factors. To bridge this gap, sampling data from large-scale genuine user–chatbot interactions, us- ing tools such as WildTeaming (Jiang et al., 2024a), can help ensure that benchmarks authentic behavior rather than rely- ing on static assumptions. Checking the ecological validity of synthetic data against actual deployment patterns further reveals critical gaps. While “in-the-wild” data collection faces privacy and transparency constraints, documenting these trade-offs increases clarity for practitioners. Each layer of abstraction weakens validity. Standard evalua- tion relies on top-down labels that often obscure the actual mechanics of risk. For instance, Ghosh et al. (2025) shows that prompts under a single label, such as “violent wrongdo- ing,” split into distinct functional clusters like operational planning versus narrative role-play. Furthermore, certain benign prompts can elicit harmful responses and cluster with known unsafe queries. This suggests risk is determined by contextual function rather than surface taxonomy. Mov- ing beyond static labels toward exploratory approaches (e.g. data-driven clustering of model output patterns) enables the discovery of granular risk categories and reveals unmapped hazards that predefined benchmarks overlook. R10. Iterative Refinement via Community Input. Safety requirements cannot be fully specified in advance. Anticipat- ing all contingencies is not possible, nor can value priorities be meaningfully articulated in the abstract, independent of concrete policy or system design choices, as emphasized in Lindblom’s insight (Narayanan, 2026). This suggests treat- ing benchmarks as evolving instruments subject to continu- ous calibration through repeated observation and revision. Recent work demonstrates this iterative approach: some benchmarks continuously calibrate using current data from news and forums (Zhang et al., 2025a), while others conduct recurring bimonthly human evaluations (Peng et al., 2021). Beyond temporal updates, involving affected communities to assess whether scenarios reflect harms they actually ex- perience can surface risks invisible to benchmark designers, reducing the distance between proxies and real-world im- pacts. When demographic groups systematically disagree on harm ratings for identical scenarios (Ali et al., 2025), this disagreement is signal, not noise. It reveals whose values current operationalizations privilege, detailed in App. C.3. 8. Case Study To demonstrate how our recommendations apply in prac- tice, we examine AIR 2024 (Zeng et al., 2024), a recent safety benchmark that aims to bridge the gap between aca- demic evaluation and real-world regulatory requirements. We assess the benchmark against three categories: construct coverage and blind spot documentation, risk quantification, and linkage to real-world deployment. Tab. 3 summarizes the assessment against our proposed checklist. Construct Coverage and Blind Spot Documentation. AIR 2024 excels in documenting coverage relative to prior benchmarks, mapping three alternatives against its taxon- omy and showing the most comprehensive covers only 71% of level-3 regulatory risk categories, with key omissions in- cluding automated decision-making, democratic deterrence, and discrimination against protected characteristics. However, documentation of blind spots beyond regulatory sources remains limited. The benchmark acknowledges its static nature, noting that risk categories require periodic up- dates, but does not specify what risks might emerge outside institutional frameworks. AIR 2024 evaluates models in isolation through single-turn interactions without examining how safety properties emerge from interactions between models, users, and deployment environments. Following our recommendation to document known blind spots (R1), articulating assumptions about deployment context and iden- tifying risk types out of scope could further strengthen va- lidity. Meanwhile, although taxonomy updates are planned, infrastructure for community input, prompt evolution, or continuous red-teaming is not yet established. Expanding evaluations to capture cumulative manipulation and context- dependent harms could strengthen practical relevance. In- corporating dynamic discovery mechanisms (R2) and re- framing known ML phenomena as safety concerns (R3) may help surface emergent risks over time. Risk Quantification. AIR 2024 uses a three-level scoring system (0, 0.5, 1) to represent harmful compliance, ambigu- ous response, and refusal, improving over binary classifica- tion by capturing intermediate outcomes. The main metric is refusal rate, defined as the percentage of responses scor- ing 1. Evaluator uncertainty is addressed through human validation of the LLM judge with Cohen’s kappa of 0.86. Several quantification limitations remain. Refusal rates are reported as point estimates without confidence intervals, treating the 89% refusal rate as definitive, rather than as a sample-based frequency. Framing these as empirical es- timates with uncertainty bounds and weighting them by real-world prevalence as shown in App. C.1 would pro- vide a more nuanced interpretation (R4). Harm severity is addressed at the taxonomy level, aligned with EU AI Act tiers (minimal, limited, high, unacceptable). Further 7 How Should AI Safety Benchmarks Benchmark Safety? consideration of how risk propagates and clarification of how scoring scales aggregate with domain-specific severity could strengthen interpretation (R5; see App. C.2 for an illustrative approach). Additionally, the benchmark does not discuss how this translates to deployment risk. Applying ex- plicit safety multipliers when extrapolating from benchmark to deployment could strengthen the actionability of reported scores (R6). Linkage to Real-World Deployment. AIR 2024 aligns benchmark results with real-world regulatory compliance. Its taxonomy is grounded in 8 government regulations and 16 corporate policies, aligning with R7’s emphasis on clar- ifying whose values define harm. Case studies mapping model performance to the EU AI Act, U.S. regulations, and corporate policies illustrate how results can guide deploy- ment decisions. Reproducibility documentation could be stronger. While GPT-4o is specified as the evaluation judge, details such as interaction mode (API vs. UI) and inference parameters are not fully recorded. Recording these details verbatim would support consistent assessments amid model updates (R8). Proxies could also be better grounded in deployment. Most prompts are LLM-generated with human review. In- corporating in-the-wild sources such as WildTeaming (Jiang et al., 2024a) could improve external validity (R9). Finally, the benchmark plans for taxonomy updates, mechanisms for continuous calibration against evolving user behavior are not established. Incorporating feedback cycles and en- gaging affected communities could reduce the gap between benchmark proxies and real-world impacts (R10). 9. Discussion Our analysis reveals that contemporary AI safety bench- marks provide an inadequate basis for asserting deployment safety. These tools offer narrow insights into specific, prede- fined behaviors of isolated models, yet struggle to capture the complex, uncertain, and socially embedded nature of safety. Consequently, strong benchmark performance can foster a false sense of security, distracting from systemic risks and perpetuating biases when benchmarks fail to ac- count for the breadth of human experience. Fairness frame- works often succeeded by simultaneously meeting the needs of scholars, businesses, advocates, and media, but confined discourse to narrow technical terms, missing fundamental issues of justice (Narayanan, 2026). Safety benchmarking risks the same: legible metrics satisfy multiple stakeholders while neglecting what matters most to affected communities. Our framework addresses this through three dimensions: construct coverage, risk quantification, and measurement validity. A potential tension is whether benchmarks should attempt to capture unknowns. We argue this is essential: capability benchmarks routinely extend boundaries (Phan et al., 2025), but safety benchmarks are far more critical, as undiscovered failures carry real-world harm. While enu- merating the unenumerable is impossible, maintaining epis- temic humility, reconsidering phenomena previously out- side safety’s scope, and investing in open-ended exploratory methods can surface risks that confirmatory testing misses. Moving toward meaningful safety evaluation suggests a shift to system-level assessment. Rather than optimizing proxies in isolation, researchers can move beyond model-centric evaluation by incorporating environmental interactions, hu- man behavioral factors, calibration with deployment data, and qualitative research with affected communities. Exam- ining AI within its sociotechnical context, with methodolo- gies that account for emergent properties and monitor for risks escaping predefined protocols, represents a promising direction. These resource-intensive approaches can poten- tially exacerbate inequalities within the research community. When comprehensive assessment is infeasible, order-of- magnitude reasoning can situate benchmarks within broader risk-reasoning workflows, and explicitly conditioning scope on deployment context and clarifying what evaluations do and do not cover offers a practical path forward. Future work should develop domain-specific treatments rec- ognizing that AI safety subcategories differ in epistemic structure and require tailored methodologies. Operationaliz- ing efficient, iterative community involvement in benchmark design also warrants investigation. 10. Conclusion AI safety is a critical concern as AI capabilities advance. While AI safety benchmarks have emerged as a popular tool for evaluation, our analysis, drawing on extensive literature, shows that they provide an incomplete and unreliable basis for assessing deployment safety. They suffer from signifi- cant gaps in scientific rigor, engineering design principles, and sociotechnical considerations. Current benchmarks are limited in their coverage of risks, fail to probabilistically quantify real-world hazards, face fundamental challenges misalign with measurement theory, and overlook that safety is embedded in complex sociotechnical systems. Effectively ensuring AI safety requires moving beyond the confines of current benchmarking practices. It necessitates devel- oping new evaluation methods and frameworks that em- brace a system-level perspective, account for uncertainty and unknown risks, are grounded in robust measurement theory, and are shaped through democratic and participatory processes that involve impacted communities. Only by ac- knowledging the inherent limitations of current technical benchmarks and adopting a more holistic approach can we hope to build and deploy AI systems that are truly safe. 8 How Should AI Safety Benchmarks Benchmark Safety? Impact Statement This paper proposes a framework for improving AI safety benchmarking by integrating principles from risk engineer- ing, measurement theory, and sociotechnical systems think- ing. By encouraging more rigorous evaluation practices, our work aims to positively influence how safety claims are validated and communicated to researchers, practitioners, and policymakers. However, proposed severity scales and safety margins risk premature standardization or co-optation for compliance theater; we address these concerns by em- phasizing iterative validation, calibration transparency, and open methodology. This work is purely methodological and releases no artifacts posing direct misuse risks. Acknowledgments We thank Dora Zhao, Jan Batzner, Shalaleh Rismani, Simon Jarvers, Naira Paola Arnez Jordan, and the I2SC Lecture Series for their valuable feedback and support. References Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Chris- tiano, John Schulman, and Dan Mané. Concrete problems in AI safety.arXivpreprintarXiv:1606.06565, 2016. Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, Pe- ter Eckersley, Ben Garfinkel, Allan Dafoe, Paul Scharre, Thomas Zeitzoff, Bobby Filar, Hyrum Anderson, et al. The malicious use of artificial intelligence: Forecasting, prevention, and mitigation. arXiv:1802.07228, 2018. URL https://arxiv.org/abs/1802.07228. Laura Weidinger, Jonathan Uesato, Jack Rae, et al. Taxonomy of risks posed by language models.In Proceedingsofthe2022ACMConferenceonFairness, Accountability,andTransparency(FAccT), 2022. doi: 10.1145/3531146.3533088. URLhttps://dl.acm. org/doi/10.1145/3531146.3533088. Joy Buolamwini and Timnit Gebru. Gender shades: In- tersectional accuracy disparities in commercial gender classification. InProceedingsofthe1stConferenceon Fairness,AccountabilityandTransparency(FAccT), vol- ume 81 ofProceedingsofMachineLearningResearch, pages 77–91, 2018. Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. RealToxicityPrompts: Eval- uating neural toxic degeneration in language mod- els.In Trevor Cohn, Yulan He, and Yang Liu, ed- itors,FindingsoftheAssociationforComputational Linguistics:EMNLP2020, pages 3356–3369, On- line, November 2020. Association for Computational Linguistics.doi: 10.18653/v1/2020.findings-emnlp. 301. URLhttps://aclanthology.org/2020. findings-emnlp.301/. NicholasCarlini,FlorianTramer,EricWallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ul- far Erlingsson, et al.Extracting training data from large language models.In30thUSENIXSecurity Symposium, 2021.URLhttps://w.usenix. org/conference/usenixsecurity21/ presentation/carlini-extracting. Fabio Urbina, Filippa Lentzos, César Invernizzi, and Sean Ekins. Dual use of artificial-intelligence-powered drug discovery.NatureMachineIntelligence, 4(3):189–191, 2022. doi: 10.1038/s42256-022-00465-9. Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard L. Phillips, Irena Gao, et al. Wilds: A benchmark of in-the-wild dis- tribution shifts. InProceedingsofthe38thInternational ConferenceonMachineLearning(ICML), volume 139 ofProceedingsofMachineLearningResearch, 2021. URLhttps://proceedings.mlr.press/ v139/koh21a.html. Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, et al. Holistic evaluation of language models.arXiv preprintarXiv:2211.09110, 2022.URLhttps:// arxiv.org/abs/2211.09110. Bertie Vidgen, Adarsh Agrawal, Ahmed M Ahmed, Victor Akinwande, Namir Al-Nuaimi, Najla Alfaraj, Elie Al- hajjar, Lora Aroyo, Trupti Bavalatti, Max Bartolo, et al. Introducing v0. 5 of the ai safety benchmark from mlcom- mons.arXivpreprintarXiv:2404.12241, 2024. Center for AI Safety. Safebench winners.https://w. mlsafety.org/safebench/winners , 2025. Ac- cessed: 2025-10-16. Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexan- der C. Berg, and Li Fei-Fei.Imagenet large scale visual recognition challenge.InternationalJournalof ComputerVision, 115(3):211–252, 2015. doi: 10.1007/ s11263-015-0816-y. Standard Performance Evaluation Corporation.Spec cpu® 2017 benchmark.https://w.spec.org/ cpu2017/, 2017. Accessed 2025-10-09. Peter Mattson, Christine Cheng, Cody Coleman, Greg Diamos, Paulius Micikevicius, David Patterson, et al. Mlperf training benchmark. InProceedingsofMachine 9 How Should AI Safety Benchmarks Benchmark Safety? LearningandSystems(MLSys), volume 2, pages 336– 349, 2020. URLhttps://proceedings.mlsys. org/paper_files/paper/2020/hash/ 411e39b117e885341f25efb8912945f7-Abstract. html. Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. Beyond accuracy: Behavioral testing of nlp models with checklist. InProceedingsofthe58th AnnualMeetingoftheAssociationforComputational Linguistics, pages 4902–4912, 2020. doi: 10.18653/v1/ 2020.acl-main.442. URLhttps://aclanthology. org/2020.acl-main.442/. Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models.CoRR, abs/2307.15043, 2023. doi: 10.48550/ARXIV.2307.15043. URLhttps://doi. org/10.48550/arXiv.2307.15043. Jacqueline Harding and Cameron Domenico Kirk-Giannini. What is ai safety?what do we want it to be? PhilosophicalStudies, pages 1–24, 2025a. Functional safety of electrical/electronic/programmable electronic safety-related systems, 2010. Nancy G. Leveson.EngineeringaSaferWorld:Systems ThinkingAppliedtoSafety. MIT Press, Cambridge, MA, 2011a. URLhttps://sunnyday.mit.edu/ safer-world.pdf. Erik Hollnagel.Safety-IandSafety-I:ThePastandFuture ofSafetyManagement. CRC Press, 2014. doi: 10.1201/ 9781315607511. ISO.Iso 26262: Road vehicles — functional safety, 2018. URLhttps://w.iso.org/standard/ 68383.html. ISO. Iso/iec 23894:2023 — artificial intelligence — guid- ance on risk management, 2023. URLhttps://w. iso.org/standard/77304.html. NIST. Artificial intelligence risk management framework (ai rmf 1.0). Technical Report NIST AI 100-1, Na- tional Institute of Standards and Technology, Gaithers- burg, MD, 2023. URLhttps://nvlpubs.nist. gov/nistpubs/ai/nist.ai.100-1.pdf. Moritz Hardt and Benjamin Recht. Patterns, predictions, and actions: A story about machine learning.arXivpreprint arXiv:2102.05242, 2021. Inioluwa Deborah Raji, Emily M Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna. Ai and the everything in the whole wide world benchmark.arXiv preprintarXiv:2111.15366, 2021. Will Orr and Edward B Kang. Ai as a sport: On the compet- itive epistemologies of benchmarking. InProceedingsof the2024ACMConferenceonFairness,Accountability, andTransparency, pages 1875–1884, 2024. Roel Dobbe. System safety and artificial intelligence. In Proceedingsofthe2022ACMConferenceonFairness, Accountability,andTransparency, pages 1584–1584, 2022. IPCC. Managing the risks of extreme events and disasters to advance climate change adaptation (srex). Technical report, Intergovernmental Panel on Climate Change, 2012. URL https://w.ipcc.ch/report/srex/. Risk management—guidelines, 2018. URLhttps:// w.iso.org/standard/65694.html. Nancy G. Leveson.EngineeringaSaferWorld:Systems ThinkingAppliedtoSafety. MIT Press, Cambridge, MA, 2011b. Homayoon Dezfuli, Allan Benjamin, Christopher Everett, Gaspare Maggio, Michael Stamatelatos, Robert Young- blood, Sergio Guarro, Peter Rutledge, James Sherrard, Curtis Smith, et al. Nasa risk management handbook. Technical report, 2011. ISO 14971:2019.Medical devices — application of risk management to medical devices. Standard ISO 14971:2019, International Organization for Standardiza- tion, Geneva, CH, 2019. European Union Agency for Railways. Final report – risk acceptance criteria for technical systems and operational procedures. Technical report, European Union Agency for Railways, 2022.URLhttps://w.era. europa.eu/system/files?file=2022-11/ risk-acceptance-criteria-for-technical-systems_ en.pdf. Accessed: 2026-01-05. DSIT.Internationalaisafetyreport2025. GOV.UK,2025.URLhttps://w. gov.uk/government/publications/ international-ai-safety-report-2025. Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. pages 3214–3252, 2022. Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Hanlin Zhang, Scott Emmons, and Dan Hendrycks. Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark. InProceedingsof the40thInternationalConferenceonMachineLearning, ICML’23. JMLR.org, 2023. 10 How Should AI Safety Benchmarks Benchmark Safety? Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zi- fan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. HarmBench: A standardized evaluation framework for automated red teaming and robust refusal.arXivpreprint arXiv:2402.04249, 2024. Dora Zhao, Jerone T. A. Andrews, Orestis Papakyriakopou- los, and Alice Xiang. Position: Measure dataset diversity, don’t just claim it. InForty-firstInternationalConference onMachineLearning,ICML2024,Vienna,Austria,July 21-27,2024. OpenReview.net, 2024a. URLhttps: //openreview.net/forum?id=jsKr6RVDDs. A. M. Bean, R. O. Kearns, A. Romanou, F. S. Hafner, H. Mayne, ..., and A. Mahdi. Measuring what matters: Construct validity in large language model benchmarks. InNeurIPS2025DatasetsandBenchmarksTrack, 2025. OpenReview id: mdA5lVvNcU. Abigail Z. Jacobs and Hanna Wallach. Measurement and fairness. InProceedingsofthe2021ACMConference onFairness,Accountability,andTransparency(FAccT), pages 375–385, 2021. doi: 10.1145/3442188.3445901. Charles A. E. Goodhart. Problems of monetary management: The uk experience. InPapersinMonetaryEconomics. Reserve Bank of Australia, 1975. Donald T. Campbell. Assessing the impact of planned so- cial change.EvaluationandProgramPlanning, 2:67–90, 1979. Anna Katariina Wisakanto, Joe Rogero, Avyay M Casheekar, and Richard Mallah. Adapting probabilistic risk assessment for ai.arXivpreprintarXiv:2504.18536, 2025. Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing. Gpt- fuzzer: Red teaming large language models with auto- generated jailbreak prompts, 2024a. URLhttps:// arxiv.org/abs/2309.10253. Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, et al. Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.AdvancesinNeuralInformation ProcessingSystems, 37:47094–47165, 2024a. Angelos Filos, Panagiotis Tigas, Rowan McAllister, Nicholas Rhinehart, Sergey Levine, and Yarin Gal. Can autonomous vehicles identify, recover from, and adapt to distribution shifts?, 2020. URLhttps://arxiv. org/abs/2006.14911. Kaichen Zhang, Bo Li, Peiyuan Zhang, Fanyi Pu, Joshua Adrian Cahyono, Kairui Hu, Shuai Liu, Yuan- han Zhang, Jingkang Yang, Chunyuan Li, and Ziwei Liu. Lmms-eval: Reality check on the evaluation of large mul- timodal models. InNAACL(Findings), pages 881–916. Association for Computational Linguistics, 2025a. Joseph Goodier and Neill D. F. Campbell. Likelihood-based out-of-distribution detection with denoising diffusion probabilistic models, 2023. URLhttps://arxiv. org/abs/2310.17432. Gabrielle Berman and Kerry Albright. Children and the data cycle: Rights and ethics in a big data world.CoRR, abs/1710.06881, 2017. URLhttp://arxiv.org/ abs/1710.06881. Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Cather- ine Olsson, Sandipan Kundu, Saurav Kadavath, et al. Discovering language model behaviors with model- written evaluations. InFindingsoftheassociationfor computationallinguistics:ACL2023, pages 13387– 13434, 2023. Junzhe Chen, Xuming Hu, Shuodi Liu, Shiyu Huang, Wei-Wei Tu, Zhaofeng He, and Lijie Wen.L- MArena: Assessing capabilities of large language mod- els in dynamic multi-agent environments.In Lun- Wei Ku, Andre Martins, and Vivek Srikumar, edi- tors,Proceedingsofthe62ndAnnualMeetingofthe AssociationforComputationalLinguistics(Volume1: LongPapers), pages 13055–13077, Bangkok, Thailand, August 2024a. Association for Computational Linguis- tics. doi: 10.18653/v1/2024.acl-long.705. URLhttps: //aclanthology.org/2024.acl-long.705/. Nassim Nicholas Taleb.TheBlackSwan:TheImpactof theHighlyImprobable. Random House, New York, 2007. ISBN 978-1400063512. Vipul Gupta, Pranav Narayanan Venkit, Hugo Laurençon, Shomir Wilson, and Rebecca J. Passonneau. Calm : A multi-task benchmark for comprehensive assessment of language model bias, 2024a. URLhttps://arxiv. org/abs/2308.12539. Mike Laszkiewicz, Imant Daunhawer, Julia E. Vogt, Asja Fischer, and Johannes Lederer. Benchmarking the fairness of image upsampling methods.InThe 2024ACMConferenceonFairness,Accountability,and Transparency,FAccT2024,RiodeJaneiro,Brazil,June 3-6,2024, pages 489–517. ACM, 2024.doi: 10. 1145/3630106.3658921. URLhttps://doi.org/ 10.1145/3630106.3658921. 11 How Should AI Safety Benchmarks Benchmark Safety? Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, et al. A strongreject for empty jailbreaks.AdvancesinNeuralInformation ProcessingSystems, 37:125416–125440, 2024. Erxin Yu, Jing Li, Ming Liao, Siqi Wang, Gao Zuchen, Fei Mi, and Lanqing Hong.CoSafe: Evaluating large language model safety in multi-turn dialogue coreference.In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Proceedingsofthe 2024ConferenceonEmpiricalMethodsinNatural LanguageProcessing, pages 17494–17508, Miami, Florida, USA, November 2024b. Association for Compu- tational Linguistics. doi: 10.18653/v1/2024.emnlp-main. 968. URLhttps://aclanthology.org/2024. emnlp-main.968/. Mi Zhang, Xudong Pan, and Min Yang. Jade: A linguistics- based safety evaluation platform for large language mod- els.arXivpreprintarXiv:2311.00286, 2023. Siyuan Wang, Zhuohan Long, Zhihao Fan, Xuan-Jing Huang, and Zhongyu Wei. Benchmark self-evolving: A multi-agent framework for dynamic llm evaluation. InProceedingsofthe31stinternationalconferenceon computationallinguistics, pages 3310–3328, 2025a. Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, et al. R-judge: Benchmarking safety risk awareness for llm agents.arXivpreprint arXiv:2401.10019, 2024. Google DeepMind. Gemini 3 pro frontier safety framework (fsf) report. Technical report, Google DeepMind, 11 2025.URLhttps://storage.googleapis. com/deepmind-media/gemini/gemini_3_ pro_fsf_report.pdf.Gemini 3 Pro Frontier Safety Framework Report. Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. CrowS-Pairs: A challenge dataset for mea- suring social biases in masked language models. In ProceedingsofEMNLP, 2020. Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta.Bold: Dataset and metrics for mea- suring biases in open-ended language generation. In Proceedingsofthe2021ACMConferenceonFairness, Accountability,andTransparency, FAccT ’21, page 862–872, New York, NY, USA, 2021. Association for Computing Machinery. ISBN 9781450383097. doi: 10. 1145/3442188.3445924. URLhttps://doi.org/ 10.1145/3442188.3445924. Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman.Bbq: A hand-built bias benchmark for question answering.arXivpreprint arXiv:2110.08193, 2021. Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li. Adversarial glue: A multi-task benchmark for robustness evaluation of language models.arXivpreprint arXiv:2111.02840, 2021. Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao. Jailbreakv: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks.InFirstConferenceonLanguage Modeling, 2024.URLhttps://openreview. net/forum?id=GC4mXVfquq. Shaona Ghosh, Heather Frase, Adina Williams, Sarah Luger, Paul Röttger, Fazl Barez, Sean McGregor, Kenneth Frick- las, Mala Kumar, Kurt Bollacker, et al. Ailuminate: Intro- ducing v1. 0 of the ai risk and reliability benchmark from mlcommons.arXivpreprintarXiv:2503.05731, 2025. Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wang- meng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. SALAD- Bench: A hierarchical and comprehensive safety bench- mark for large language models.InFindingsof theAssociationforComputationalLinguistics(ACL), 2024a. Melissa Hall, Laura Gustafson, Aaron Adcock, Ishan Misra, and Candace Ross. Vision-language models perform- ing zero-shot tasks exhibit disparities between gender groups. In2023IEEE/CVFInternationalConference onComputerVisionWorkshops(ICCVW), pages 2770– 2777, 2023. doi: 10.1109/ICCVW60793.2023.00294. Fengqing Jiang, Zhangchen Xu, Luyao Niu, Zhen Xi- ang, Bhaskar Ramasubramanian, Bo Li, and Radha Poovendran. Artprompt: Ascii art-based jailbreak at- tacks against aligned llms. InProceedingsofthe62nd AnnualMeetingoftheAssociationforComputational Linguistics(Volume1:LongPapers), pages 15157– 15173, 2024b. Jacob Dineen, Aswin RRV, Qin Liu, Zhikun Xu, Xiao Ye, Ming Shen, Zhaonan Li, Shijie Lu, Chitta Baral, Muhao Chen, et al. Qa-lign: Aligning llms through constitution- ally decomposed qa.arXivpreprintarXiv:2506.08123, 2025. Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou. Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions. arXivpreprintarXiv:2309.07875, 2023. 12 How Should AI Safety Benchmarks Benchmark Safety? Tessa Han, Aounon Kumar, Chirag Agarwal, and Himabindu Lakkaraju. Medsafetybench: Evaluating and improving the medical safety of large language models. InTheThirty-eightConferenceonNeuralInformation ProcessingSystemsDatasetsandBenchmarksTrack, 2024. URLhttps://openreview.net/forum? id=cFyagd2Yh4. Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang."do anything now":Charac- terizing and evaluating in-the-wild jailbreak prompts on large language models.InProceedingsofthe 2024onACMSIGSACConferenceonComputerand CommunicationsSecurity, CCS ’24, page 1671–1685, New York, NY, USA, 2024a. Association for Com- puting Machinery. ISBN 9798400706363. doi: 10. 1145/3658644.3670388. URLhttps://doi.org/ 10.1145/3658644.3670388. Yufei Huang and Deyi Xiong. CBBQ: A Chinese bias bench- mark dataset curated with human-AI collaboration for large language models. In Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, editors,Proceedingsofthe2024Joint InternationalConferenceonComputationalLinguistics, LanguageResourcesandEvaluation(LREC-COLING 2024), pages 2917–2929, Torino, Italia, May 2024. ELRA and ICCL. URLhttps://aclanthology.org/ 2024.lrec-main.260/. Jiaxu Zhao, Meng Fang, Zijing Shi, Yitong Li, Ling Chen, and Mykola Pechenizkiy. CHBias: Bias evaluation and mitigation of Chinese conversational language models. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors,Proceedingsofthe61stAnnualMeetingofthe AssociationforComputationalLinguistics(Volume1: LongPapers), pages 13538–13556, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.757.URLhttps: //aclanthology.org/2023.acl-long.757/. Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. Med-HALT: Medical domain hallucina- tion test for large language models. In Jing Jiang, David Reitter, and Shumin Deng, editors,Proceedingsofthe 27thConferenceonComputationalNaturalLanguage Learning(CoNLL), pages 314–334, Singapore, Decem- ber 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.conll-1.21. URLhttps:// aclanthology.org/2023.conll-1.21/. Yi Zeng, Yu Yang, Andy Zhou, Jeffrey Ziwei Tan, Yuheng Tu, Yifan Mai, Kevin Klyman, Minzhou Pan, Ruoxi Jia, Dawn Song, et al. Air-bench 2024: A safety benchmark based on risk categories from regulations and policies. arXivpreprintarXiv:2407.17436, 2024. Enrico Fermi. My observations during the explosion at trinity on july 16, 1945. Technical report, Los Alamos National Laboratory, 1945. Lu Zhang, Yongkai Wu, and Xintao Wu. Achieving non- discrimination in prediction, 2018. URLhttps:// arxiv.org/abs/1703.00060. Hyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi, and Maarten Sap. Prosocialdialog: A prosocial backbone for conversa- tional agents.arXivpreprintarXiv:2205.12688, 2022. Michael L Dourson and Jerry F Stara. Regulatory history and experimental support of uncertainty (safety) factors. Regulatorytoxicologyandpharmacology, 3(3):224–238, 1983. Eran Tal. Measurement in Science. In Edward N. Zalta, editor,TheStanfordEncyclopediaofPhilosophy. Meta- physics Research Lab, Stanford University, Fall 2020 edition, 2020. Yongshuo Zong, Yongxin Yang, and Timothy Hospedales. Medfair: Benchmarking fairness for medical imaging. InInternationalConferenceonLearningRepresentations (ICLR), 2023. Olawale Salaudeen, Anka Reuel, Ahmed Ahmed, Suhana Bedi, Zachary Robertson, Sudharsan Sundar, Ben Domingue, Angelina Wang, and Sanmi Koyejo. Mea- surement to meaning: A validity-centered framework for ai evaluation.arXivpreprintarXiv:2505.10573, 2025. Hanna Wallach, Meera Desai, A Feder Cooper, Angelina Wang, Chad Atalla, Solon Barocas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P Alex Dow, et al. Position: Evaluating generative ai systems is a social science measurement challenge.arXivpreprint arXiv:2502.00561, 2025. Ivo Petrov, Jasper Dekoninck, Lyuben Baltadzhiev, Maria Drencheva, Kristian Minchev, Mislav Balunovi ́ c, Nikola Jovanovi ́ c, and Martin Vechev. Proof or bluff? evalu- ating llms on 2025 usa math olympiad.arXivpreprint arXiv:2503.21934, 2025. Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe- bench: Can language models resolve real-world github issues?arXivpreprintarXiv:2310.06770, 2023. Andrew D Selbst, Danah Boyd, Sorelle A Friedler, Suresh Venkatasubramanian, and Janet Vertesi. Fairness and ab- straction in sociotechnical systems. InProceedingsofthe conferenceonfairness,accountability,andtransparency, pages 59–68, 2019. 13 How Should AI Safety Benchmarks Benchmark Safety? Shalaleh Rismani, Renee Shelby, Leah Davis, Negar Rostamzadeh, and AJung Moon.Measuring what matters: Connecting ai ethics evaluations to system attributes, hazards, and harms.Proceedingsofthe AAAI/ACMConferenceonAI,Ethics,andSociety, 8 (3):2199–2213, Oct. 2025.doi: 10.1609/aies.v8i3. 36706.URLhttps://ojs.aaai.org/index. php/AIES/article/view/36706. Arvind Narayanan. What if algorithmic fairness is a cate- gory error? In Sven Nyholm, Atoosa Kasirzadeh, and John Zerilli, editors,ContemporaryDebatesintheEthics ofArtificialIntelligence, pages 77–96. Wiley-Blackwell, 2026. ISBN 9781394258819. Baolin Peng, Chunyuan Li, Zhu Zhang, Chenguang Zhu, Jinchao Li, and Jianfeng Gao. RADDLE: An evaluation benchmark and analysis platform for robust task-oriented dialog systems. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors,Proceedingsofthe59th AnnualMeetingoftheAssociationforComputational Linguisticsandthe11thInternationalJointConference onNaturalLanguageProcessing(Volume1:Long Papers), pages 4418–4429, Online, August 2021. Associ- ation for Computational Linguistics. doi: 10.18653/v1/ 2021.acl-long.341. URLhttps://aclanthology. org/2021.acl-long.341/. Dalia Ali, Dora Zhao, Allison Koenecke, and Orestis Pa- pakyriakopoulos. Operationalizing pluralistic values in large language model alignment reveals trade-offs in safety, inclusivity, and model behavior.arXivpreprint arXiv:2511.14476, 2025. Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al. Humanity’s last exam.arXivpreprintarXiv:2501.14249, 2025. Jacqueline Harding and Cameron Domenico Kirk-Giannini. What is ai safety?what do we want it to be? PhilosophicalStudies, 182:1495–1518, 2025b.doi: 10.1007/s11098-025-02367-z. Greg Guest, Arwen Bunce, and Laura Johnson. How many interviews are enough? an experiment with data satura- tion and variability.Fieldmethods, 18(1):59–82, 2006. Barney Glaser and Anselm Strauss.Discoveryofgrounded theory:Strategiesforqualitativeresearch. Routledge, 2017. Muhammad Naeem, Wilson Ozuem, Kerry Howell, and Silvia Ranfagni. Demystification and actualisation of data saturation in qualitative research through thematic analysis.InternationalJournalofQualitativeMethods, 23:16094069241229777, 2024. Benjamin Saunders, Julius Sim, Tom Kingstone, Shula Baker, Jackie Waterfield, Bernadette Bartlam, Heather Burroughs, and Clare Jinks. Saturation in qualitative research: exploring its conceptualization and operational- ization.Quality&quantity, 52(4):1893–1907, 2018. Monique Hennink and Bonnie N Kaiser. Sample sizes for saturation in qualitative research: A systematic review of empirical tests.Socialscience&medicine, 292:114523, 2022. Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. Toxigen: A large-scale machine-generated dataset for adversar- ial and implicit hate speech detection.arXivpreprint arXiv:2203.09509, 2022. Aakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant, Julia Kreutzer, Marzieh Fadaee, and Sara Hooker. The multilingual alignment prism: Aligning global and local preferences to reduce harm. InEMNLP, pages 12027–12049. Association for Computational Lin- guistics, 2024. Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. Fine-tuning aligned language models compromises safety, even when users do not intend to!InTheTwelfthInternational ConferenceonLearningRepresentations,ICLR2024, Vienna,Austria,May7-11,2024. OpenReview.net, 2024. URLhttps://openreview.net/forum? id=hTEGyKf0dZ. Richard Ren, Steven Basart, Adam Khoja, Alice Gatti, Long Phan, Xuwang Yin, Mantas Mazeika, Alexander Pan, Gabriel Mukobi, Ryan Kim, et al. Safetywashing: Do ai safety benchmarks actually measure safety progress? AdvancesinNeuralInformationProcessingSystems, 37: 68559–68594, 2024. CRFM. Air-bench leaderboard, holistic evaluation of lan- guage models (helm).https://crfm.stanford. edu/helm/air-bench/latest/ , 2024. Accessed: 2025-12-27. Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. WildChat: 1m chatgpt in- teraction logs in the wild. InProceedingsofthe12th InternationalConferenceonLearningRepresentations (ICLR), 2024b. Lawrence Weinstein and John A. Adam.Guesstimation: SolvingtheWorld’sProblemsontheBackofaCocktail Napkin. Princeton University Press, 2008. Sanjoy Mahajan.Street-FightingMathematics:TheArtof EducatedGuessingandOpportunisticProblemSolving. MIT Press, 2010. 14 How Should AI Safety Benchmarks Benchmark Safety? Hans Christian von Baeyer.TheFermiSolution:Essayson Science. Random House, 1993. RN Jones, R Leemans, LO Mearns, N Nakicenovic, AB Pittock, SM Semenov, S Gromov, SR Khan, and A Koukhta.Developing and applying scenar- ios.ClimateChange2001:Impacts,Adaptation,and Vulnerability:ContributionofWorkingGroupIItothe ThirdAssessmentReportoftheIntergovernmentalPanel onClimateChange, 2:145, 2001. Hannah Kosow and Robert Gaßner.Methodsoffutureand scenarioanalysis:overview,assessment,andselection criteria, volume 39. DEU, 2008. Stanley Kaplan and B. John Garrick. On the quantitative definition of risk.RiskAnalysis, 1(1):11–27, 1981. Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. " do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. InProceedingsofthe2024onACM SIGSACConferenceonComputerandCommunications Security, pages 1671–1685, 2024b. Transactional Records Access Clearinghouse (TRAC). Copyright infringement litigation fell 22 percent in fy 2016.https://tracreports.org/ whatsnew/email.161121.html, 2016.URL https://tracreports.org/whatsnew/ email.161121.html.Report noting 3,944 new federal copyright infringement cases filed in FY 2016. Adam Holland,Christopher Bavitz,and Lumen Database project.European commission stake- holder dialogue on article 17 – written statement (lumen).https://communia-association. org/wp-content/uploads/2019/12/ stakholderdiagogue4_Lumen.pdf ,2019. URLhttps://communia-association. org/wp-content/uploads/2019/12/ stakholderdiagogue4_Lumen.pdf.As of Dec. 2019, Lumen database contains approximately 12 million takedown notices, majority DMCA (copyright) notices. NerdyNav. ChatGPT statistics 2025: 800m+ users, rev- enue, and key insights.https://nerdynav.com/ chatgpt-statistics/ , 2025. Accessed: January 2026. Benjamin Brady, Roy Germano, and Christopher Sprigman. Dataset of statutory damages awards in u.s. copyright cases (2009–2020), 2025. URLhttps://doi.org/ 10.58153/jyn8d-n9f12 . Contains 277 statutory damages awards from 2009–2020. Chloé Touzet, Henry Papadatos, Malcolm Murray, Otter Quarks, Steve Barrett, Alejandro Tlaie Boria, Elija Per- rier, Matthew Smith, and Siméon Campos. The role of risk modeling in advanced ai risk management.arXiv preprintarXiv:2512.08723, 2025. Sean McGregor, Victor Lu, Vassil Tashev, Armstrong Found- jem, Aishwarya Ramasethu, Sadegh AlMahdi Kazemi Zarkouei, Chris Knotz, Kongtao Chen, Alicia Parrish, Anka Reuel, et al. Risk management for mitigating benchmark failure modes: Benchrisk.arXivpreprint arXiv:2510.21460, 2025. F Niehaus. Use of probabilistic safety assessment (psa) for nuclear installations.Safetyscience, 40(1-4):153–176, 2002. George E Apostolakis. How useful is quantitative risk as- sessment?RiskAnalysis:AnInternationalJournal, 24 (3):515–520, 2004. Enrico Zio. Reliability engineering: Old problems and new challenges.Reliabilityengineering&systemsafety, 94 (2):125–141, 2009. Lawrence J Hettinger, Alex Kirlik, Yang Miang Goh, and Peter Buckle. Modelling and simulation of complex so- ciotechnical systems: envisioning and analysing work environments.Ergonomics, 58(4):600–614, 2015. Sacit M Cetiner, Michael David Muhlheim, George F Flana- gan, David L Fugate, and Roger A Kisner. Development of an automated decision-making tool for supervisory control system. Technical report, Oak Ridge National Lab.(ORNL), Oak Ridge, TN (United States), 2014. Jamy Li and Mark Chignell. Fmea-ai: Ai fairness impact assessment using failure mode and effects analysis.AI andEthics, 2(4):837–850, 2022. Simon Mylius. Systematic hazard analysis for frontier ai using stpa.arXivpreprintarXiv:2506.01782, 2025. Rishabh Bhardwaj, Duc Anh Do, and Soujanya Poria. Lan- guage models are homer simpson! safety re-alignment of fine-tuned language models through task arithmetic. InProceedingsofthe62ndAnnualMeetingofthe AssociationforComputationalLinguistics(Volume1: LongPapers), pages 14138–14149, 2024. Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams. " i’m sorry to hear that": Finding new biases in language models with a holis- tic descriptor dataset.arXivpreprintarXiv:2205.09209, 2022. Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia Chilton, Desmond Patton, Kathleen McKeown, and 15 How Should AI Safety Benchmarks Benchmark Safety? William Yang Wang. Safetext: A benchmark for explor- ing physical safety in language models.arXivpreprint arXiv:2210.10045, 2022. Prannaya Gupta, Le Qi Yau, Hao Han Low, I-Shiang Lee, Hugo Maximus Lim, Yu Xin Teoh, Koh Jia Hng, Dar Win Liew, Rishabh Bhardwaj, Rajat Bhardwaj, et al. Wallede- val: A comprehensive safety evaluation toolkit for large language models. InProceedingsofthe2024Conference onEmpiricalMethodsinNaturalLanguageProcessing: SystemDemonstrations, pages 397–407, 2024b. Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. Xstest: A test suite for identifying exaggerated safety behaviours in large language models.arXivpreprintarXiv:2308.01263, 2023. Simon Malberg, Roman Poletukhin, Carolin M Schuster, and Georg Groh. A comprehensive evaluation of cog- nitive biases in llms.arXivpreprintarXiv:2410.15413, 2024. Ashutosh Sathe, Prachi Jain, and Sunayana Sitaram. A unified framework and dataset for assessing soci- etal bias in vision-language models.arXivpreprint arXiv:2402.13636, 2024. Bhaktipriya Radharapu, Kevin Robinson, Lora Aroyo, and Preethi Lahoti. Aart: Ai-assisted red-teaming with diverse data generation for new llm-powered applications.arXiv preprintarXiv:2311.08592, 2023. Abhay Gupta, Philip Meng, Ece Yurtseven, Sean O’Brien, and Kevin Zhu. Aavenue: Detecting llm biases on nlu tasks in aave via a novel benchmark.arXivpreprint arXiv:2408.14845, 2024c. Jianwu Fang, Lei-lei Li, Junfei Zhou, Junbin Xiao, Hongkai Yu, Chen Lv, Jianru Xue, and Tat-Seng Chua. Abductive ego-view accident video understanding for safe driving perception. InProceedingsoftheIEEE/CVFConference onComputerVisionandPatternRecognition, pages 22030–22040, 2024. Linjie Li, Jie Lei, Zhe Gan, and Jingjing Liu. Adversar- ial vqa: A new benchmark for evaluating the robust- ness of vqa models. InProceedingsoftheIEEE/CVF InternationalConferenceonComputerVision, pages 2042–2051, 2021. Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr. Agent- dojo: A dynamic environment to evaluate prompt in- jection attacks and defenses for llm agents.Advances inNeuralInformationProcessingSystems, 37:82895– 82920, 2024. Simone Tedeschi, Felix Friedrich, Patrick Schramowski, Kristian Kersting, Roberto Navigli, Huu Nguyen, and Bo Li. Alert: A comprehensive benchmark for assessing large language models’ safety through red teaming.arXiv preprintarXiv:2404.08676, 2024. Wenxuan Wang, Zhaopeng Tu, Chang Chen, Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, and Michael Lyu. All languages matter: On the multilingual safety of llms. InFindingsoftheAssociationforComputational Linguistics:ACL2024, pages 5865–5877, 2024a. Anna Sotnikova, Yang Trista Cao, Hal Daumé I, and Rachel Rudinger. Analyzing stereotypes in generative text inference tasks. InFindingsoftheAssociationfor ComputationalLinguistics:ACL-IJCNLP2021, pages 4052–4065, 2021. Mahmoud Alfadel, Diego Elias Costa, Mouafak Mokhal- lalati, Emad Shihab, and Bram Adams. On the threat of npm vulnerable dependencies in node. js applications. arXivpreprintarXiv:2009.09019, 2020. Vikram Nitin, Rahul Krishna, Luiz Lemos do Valle, and Baishakhi Ray. C2saferrust: Transforming c projects into safer rust with neurosymbolic techniques.arXivpreprint arXiv:2501.14257, 2025. Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang. Are large pre-trained language models leaking your per- sonal information?arXivpreprintarXiv:2205.12628, 2022. Yixin Wan, Jieyu Zhao, Aman Chadha, Nanyun Peng, and Kai-Wei Chang. Are personalized stochastic parrots more dangerous? evaluating persona biases in dialogue sys- tems.arXivpreprintarXiv:2310.05280, 2023. Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen. Catastrophic jailbreak of open- source llms via exploiting generation.arXivpreprint arXiv:2310.06987, 2023. Tanmana Sadhu, Ali Pesaranghader, Yanan Chen, and Dong Hoon Yi.Athena: Safe autonomous agents with verbal contrastive learning.arXivpreprint arXiv:2408.11021, 2024. Yige Li, Hanxun Huang, Yunhan Zhao, Xingjun Ma, and Jun Sun. Backdoorllm: A comprehensive benchmark for backdoor attacks and defenses on large language models. arXivpreprintarXiv:2408.12798, 2024b. Jiho Jin, Woosung Kang, Junho Myung, and Alice Oh. Social bias benchmark for generation: A comparison of generation and qa-based evaluations.arXivpreprint arXiv:2503.06987, 2025a. 16 How Should AI Safety Benchmarks Benchmark Safety? Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. Beavertails: Towards improved safety alignment of llm via a human-preference dataset. AdvancesinNeuralInformationProcessingSystems, 36: 24678–24704, 2023. Yoonshik Kim and Jaeyoon Jung.Koffvqa:An objectively evaluated free-form vqa benchmark for large vision-language models in the korean language. InProceedingsoftheComputerVisionandPattern RecognitionConference, pages 575–585, 2025. Kexin Huang, Xiangyang Liu, Qianyu Guo, Tianxiang Sun, Jiawei Sun, Yaru Wang, Zeyang Zhou, Yixu Wang, Yan Teng, Xipeng Qiu, et al. Flames: Benchmarking value alignment of llms in chinese. InProceedingsof the2024ConferenceoftheNorthAmericanChapterof theAssociationforComputationalLinguistics:Human LanguageTechnologies(Volume1:LongPapers), pages 4551–4591, 2024a. Hao Liang, Pietro Perona, and Guha Balakrishnan. Bench- marking algorithmic bias in face recognition: An experi- mental approach using synthetic faces and human eval- uation. In2023IEEE/CVFInternationalConferenceon ComputerVision(ICCV), pages 4954–4964, 2023. doi: 10.1109/ICCV51070.2023.00459. Nikhil Verma and Manasa Bharadwaj. The hidden space of safety: Understanding preference-tuned llms in multilin- gual context.arXivpreprintarXiv:2504.02708, 2025. Jia Yu, Long Li, and Zhenzhong Lan. Beyond binary clas- sification: A fine-grained safety dataset for large lan- guage models.IEEEAccess, 12:64717–64726, 2024c. doi: 10.1109/ACCESS.2024.3393245. URLhttps: //doi.org/10.1109/ACCESS.2024.3393245. Hannah Kirk, Yennie Jun, Haider Iqbal, Elias Benussi, Fil- ippo Volpin, Frederic A. Dreyer, Aleksandar Shtedrit- ski, and Yuki M. Asano.Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models, 2021. URL https://arxiv.org/abs/2102.04130. Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Ja- son Weston, and Emily Dinan.Bot-adversarial di- alogue for safe conversational agents.In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cot- terell, Tanmoy Chakraborty, and Yichao Zhou, edi- tors,Proceedingsofthe2021ConferenceoftheNorth AmericanChapteroftheAssociationforComputational Linguistics:HumanLanguageTechnologies, pages 2950–2968, Online, June 2021. Association for Compu- tational Linguistics. doi: 10.18653/v1/2021.naacl-main. 235. URLhttps://aclanthology.org/2021. naacl-main.235/. Siddharth D Jaiswal, Animesh Ganai, Abhisek Dash, Sap- tarshi Ghosh, and Animesh Mukherjee. Breaking the global north stereotype: A global south-centric bench- mark dataset for auditing and mitigating biases in facial recognition systems, 2024. URLhttps://arxiv. org/abs/2407.15810. Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. Build it break it fix it for dialogue safety: Robustness from adversarial human attack. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, edi- tors,Proceedingsofthe2019ConferenceonEmpirical MethodsinNaturalLanguageProcessingandthe9th InternationalJointConferenceonNaturalLanguage Processing(EMNLP-IJCNLP), pages 4537–4546, Hong Kong, China, November 2019. Association for Compu- tational Linguistics. doi: 10.18653/v1/D19-1461. URL https://aclanthology.org/D19-1461/. Edvard P. Bjørgen, Simen Madsen, Therese S. Bjørk- nes, Fredrik V. Heimsæter, Robin Håvik, Morten Lin- derud, Per-Niklas Longberg, Louise A. Dennis, and Marija Slavkovik. Cake, death, and trolleys: Dilem- mas as benchmarks of ethical decision-making. AIES ’18, page 23–29, New York, NY, USA, 2018. Associa- tion for Computing Machinery. ISBN 9781450360128. doi: 10.1145/3278721.3278767. URLhttps://doi. org/10.1145/3278721.3278767. Yufan Chen, Arjun Arunasalam, and Z. Berkay Celik. Can large language models provide security & privacy ad- vice? measuring the ability of llms to refute miscon- ceptions. InProceedingsofthe39thAnnualComputer SecurityApplicationsConference, ACSAC ’23, page 366–378, New York, NY, USA, 2023. Association for Computing Machinery. ISBN 9798400708862. doi: 10. 1145/3627106.3627196. URLhttps://doi.org/ 10.1145/3627106.3627196. Norman Mu, Sarah Chen, Zifan Wang, Sizhe Chen, David Karamardian, Lulwa Aljeraisy, Basel Alomair, Dan Hendrycks, and David Wagner. Can llms follow sim- ple rules?, 2024. URLhttps://arxiv.org/abs/ 2311.04235. Rashidul Islam, Shimei Pan, and James R. Foulds. Can we obtain fairness for free?InProceedings ofthe2021AAAI/ACMConferenceonAI,Ethics, andSociety, AIES ’21, page 586–596, New York, NY, USA, 2021. Association for Computing Machin- ery. ISBN 9781450384735. doi: 10.1145/3461702. 3462614.URLhttps://doi.org/10.1145/ 3461702.3462614. 17 How Should AI Safety Benchmarks Benchmark Safety? Yuen Chen, Vethavikashini Chithrra Raghuram, Justus Mattern, Rada Mihalcea, and Zhijing Jin. Causally testing gender bias in LLMs: A case study on oc- cupational bias. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors,FindingsoftheAssociationfor ComputationalLinguistics:NAACL2025, pages 4984– 5004, Albuquerque, New Mexico, April 2025. Asso- ciation for Computational Linguistics. ISBN 979-8- 89176-195-7.doi: 10.18653/v1/2025.findings-naacl. 281. URLhttps://aclanthology.org/2025. findings-naacl.281/. Yuhang Wang, Yanxu Zhu, Chao Kong, Shuyu Wei, Xi- aoyuan Yi, Xing Xie, and Jitao Sang.CDEval: A benchmark for measuring the cultural dimensions of large language models. In Vinodkumar Prabhakaran, Sunipa Dev, Luciana Benotti, Daniel Hershcovich, Laura Cabello, Yong Cao, Ife Adebara, and Li Zhou, edi- tors,Proceedingsofthe2ndWorkshoponCross-Cultural ConsiderationsinNLP, pages 1–16, Bangkok, Thailand, August 2024b. Association for Computational Linguis- tics. doi: 10.18653/v1/2024.c3nlp-1.1. URLhttps: //aclanthology.org/2024.c3nlp-1.1/. Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu, Meijuan An, Bikun Yang, KaiKai Zhao, Kai Wang, and Shiguo Lian. Chisafetybench: A chinese hierarchical safety benchmark for large language models, 2024a. URL https://arxiv.org/abs/2406.10311. Yizhi Li, Ge Zhang, Xingwei Qu, Jiali Li, Zhaoqun Li, Noah Wang, Hao Li, Ruibin Yuan, Yinghao Ma, Kai Zhang, Wangchunshu Zhou, Yiming Liang, Lei Zhang, Lei Ma, Jiajun Zhang, Zuowen Li, Wenhao Huang, Chenghua Lin, and Jie Fu. CIF-bench: A Chinese instruction- following benchmark for evaluating the generalizability of large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,FindingsoftheAssociation forComputationalLinguistics:ACL2024, pages 12431– 12446, Bangkok, Thailand, August 2024c. Association for Computational Linguistics. doi: 10.18653/v1/2024. findings-acl.739. URLhttps://aclanthology. org/2024.findings-acl.739/. Haiyang Zheng, Nan Pu, Wenjing Li, Teng Long, Nicu Sebe, and Zhun Zhong. Open-world deepfake attribution via confidence-aware asymmetric learning, 2025. URL https://arxiv.org/abs/2512.12667. Yiju Guo, Ganqu Cui, Lifan Yuan, Ning Ding, Zexu Sun, Bowen Sun, Huimin Chen, Ruobing Xie, Jie Zhou, Yankai Lin, Zhiyuan Liu, and Maosong Sun.Con- trollable preference optimization: Toward controllable multi-objective alignment. In Yaser Al-Onaizan, Mo- hit Bansal, and Yun-Nung Chen, editors,Proceedings ofthe2024ConferenceonEmpiricalMethodsin NaturalLanguageProcessing, pages 1437–1454, Miami, Florida, USA, November 2024. Association for Compu- tational Linguistics. doi: 10.18653/v1/2024.emnlp-main. 85.URLhttps://aclanthology.org/2024. emnlp-main.85/. Peiran Wang, Xiaogeng Liu, and Chaowei Xiao. CVE- bench: Benchmarking LLM-based software engineer- ing agent’s ability to repair real-world CVE vulnera- bilities. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors,Proceedingsofthe2025Conferenceofthe NationsoftheAmericasChapteroftheAssociation forComputationalLinguistics:HumanLanguage Technologies(Volume1:LongPapers), pages 4207– 4224, Albuquerque, New Mexico, April 2025b. As- sociation for Computational Linguistics. ISBN 979- 8-89176-189-6.doi: 10.18653/v1/2025.naacl-long. 212. URLhttps://aclanthology.org/2025. naacl-long.212/. Andy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders, Justin W. Lin, Eliot Jones, Gashon Hussein, Samantha Liu, Donovan Jasper, Pura Peetathawatchai, Ari Glenn, Vikram Sivashankar, Daniel Zamoshchin, Leo Glikbarg, Derek Askaryar, Mike Yang, Teddy Zhang, Rishi Alluri, Nathan Tran, Rinnara Sangpisit, Polycar- pos Yiorkadjis, Kenny Osele, Gautham Raghupathi, Dan Boneh, Daniel E. Ho, and Percy Liang.Cybench: A framework for evaluating cybersecurity capabilities and risks of language models, 2025b. URLhttps: //arxiv.org/abs/2408.08926. Nithya Sambasivan, Erin Arnesen, Ben Hutchinson, Tulsee Doshi, and Vinodkumar Prabhakaran. Re-imagining algorithmic fairness in india and beyond. FAccT ’21, page 315–328, New York, NY, USA, 2021. Associa- tion for Computing Machinery. ISBN 9781450383097. doi: 10.1145/3442188.3445896. URLhttps://doi. org/10.1145/3442188.3445896. David Sun, Artem Abzaliev, Hadas Kotek, Christo- pher Klein, Zidi Xiu, and Jason Williams.DEL- PHI: Data for evaluating LLMs’ performance in handling controversial issues.In Mingxuan Wang and Imed Zitouni, editors,Proceedingsofthe2023 ConferenceonEmpiricalMethodsinNaturalLanguage Processing:IndustryTrack, pages 820–827, Singa- pore, December 2023a. Association for Computational Linguistics.doi: 10.18653/v1/2023.emnlp-industry. 76.URLhttps://aclanthology.org/2023. emnlp-industry.76/. Lora Aroyo, Alex S. Taylor, Mark Díaz, Christopher M. Homan, Alicia Parrish, Greg Serapio-García, Vinodku- mar Prabhakaran, and Ding Wang. Dices dataset: di- versity in conversational ai evaluation for safety. In 18 How Should AI Safety Benchmarks Benchmark Safety? Proceedingsofthe37thInternationalConferenceon NeuralInformationProcessingSystems, NIPS ’23, Red Hook, NY, USA, 2023. Curran Associates Inc. Alex Tamkin, Amanda Askell, Liane Lovitt, Esin Durmus, Nicholas Joseph, Shauna Kravec, Karina Nguyen, Jared Kaplan, and Deep Ganguli. Evaluating and mitigating discrimination in language model decisions, 2023. URL https://arxiv.org/abs/2312.03689. Rajat Rawat, Hudson McBride, Rajarshi Ghosh, Dhiyaan Nirmal, Jong Moon, Dhruv Alamuri, Sean O’Brien, and Kevin Zhu. DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis us- ing large language models.In Daryna Dementieva, Oana Ignat, Zhijing Jin, Rada Mihalcea, Giorgio Pi- atti, Joel Tetreault, Steven Wilson, and Jieyu Zhao, ed- itors,ProceedingsoftheThirdWorkshoponNLPfor PositiveImpact, pages 334–348, Miami, Florida, USA, November 2024. Association for Computational Linguis- tics. doi: 10.18653/v1/2024.nlp4pi-1.29. URLhttps: //aclanthology.org/2024.nlp4pi-1.29/. Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin. Do-not-answer: Evaluating safe- guards in LLMs. In Yvette Graham and Matthew Purver, editors,FindingsoftheAssociationforComputational Linguistics:EACL2024, pages 896–911, St. Julian’s, Malta, March 2024c. Association for Computational Linguistics. URLhttps://aclanthology.org/ 2024.findings-eacl.61/. Haochen Liu, Jamell Dacon, Wenqi Fan, Hui Liu, Zi- tao Liu, and Jiliang Tang. Does gender matter? to- wards fairness in dialogue systems. In Donia Scott, Nuria Bel, and Chengqing Zong, editors,Proceedings ofthe28thInternationalConferenceonComputational Linguistics, pages 4403–4416, Barcelona, Spain (Online), December 2020. International Committee on Computa- tional Linguistics. doi: 10.18653/v1/2020.coling-main. 390. URLhttps://aclanthology.org/2020. coling-main.390/. Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi. Can llms keep a secret? testing privacy implications of language models via contextual integrity theory, 2024. URL https://arxiv.org/abs/2310.17884. Tetsushi Ohki, Yuya Sato, Masakatsu Nishigaki, and Koichi Ito. Labellessface: Fair metric learning for face recog- nition without attribute labels, 2024. URLhttps: //arxiv.org/abs/2409.09274. Sayan Ghosh and Shashank Srivastava. ePiC: Employing proverbs in context as a benchmark for abstract language understanding. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors,Proceedingsofthe60th AnnualMeetingoftheAssociationforComputational Linguistics(Volume1:LongPapers), pages 3989–4004, Dublin, Ireland, May 2022. Association for Compu- tational Linguistics. doi: 10.18653/v1/2022.acl-long. 276. URLhttps://aclanthology.org/2022. acl-long.276/. Divij Bajaj, Yuanyuan Lei, Jonathan Tong, and Rui- hong Huang.Evaluating gender bias of LLMs in making morality judgements.In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Findings oftheAssociationforComputationalLinguistics: EMNLP2024, pages 15804–15818, Miami, Florida, USA, November 2024. Association for Computational Linguistics.doi: 10.18653/v1/2024.findings-emnlp. 928. URLhttps://aclanthology.org/2024. findings-emnlp.928/. Junjie Wu, Tsz Ting Chung, Kai Chen, and Dit-Yan Yeung. Unified triplet-level hallucination evaluation for large vision-language models, 2025. URLhttps://arxiv. org/abs/2410.23114. Zekun Li, Baolin Peng, Pengcheng He, and Xifeng Yan. Evaluating the instruction-following robustness of large language models to prompt injection.In Yaser Al- Onaizan, Mohit Bansal, and Yun-Nung Chen, edi- tors,Proceedingsofthe2024ConferenceonEmpirical MethodsinNaturalLanguageProcessing, pages 557– 568, Miami, Florida, USA, November 2024d. Association for Computational Linguistics. doi: 10.18653/v1/2024. emnlp-main.33.URLhttps://aclanthology. org/2024.emnlp-main.33/. Svetlana Kiritchenko and Saif Mohammad. Examining gen- der and race bias in two hundred sentiment analysis sys- tems. In Malvina Nissim, Jonathan Berant, and Alessan- dro Lenci, editors,ProceedingsoftheSeventhJoint ConferenceonLexicalandComputationalSemantics, pages 43–53, New Orleans, Louisiana, June 2018. Asso- ciation for Computational Linguistics. doi: 10.18653/v1/ S18-2005. URLhttps://aclanthology.org/ S18-2005/. Xiaoshuai Wu, Xin Liao, Bo Ou, Yuling Liu, and Zheng Qin. Are watermarks bugs for deepfake detectors? rethinking proactive forensics. InProceedingsoftheThirty-Third InternationalJointConferenceonArtificialIntelligence, 2024. Laura Gustafson, Chloe Rolland, Nikhila Ravi, Quentin Duval, Aaron Adcock, Cheng-Yang Fu, Melissa Hall, and Candace Ross. Facet: Fairness in computer vision evaluation benchmark. In2023IEEE/CVFInternational ConferenceonComputerVision(ICCV), pages 20313– 20325, 2023. doi: 10.1109/ICCV51070.2023.01863. 19 How Should AI Safety Benchmarks Benchmark Safety? Ilias Chalkidis, Tommaso Pasini, Sheng Zhang, Letizia Tomada, Sebastian Schwemer, and Anders Søgaard. Fair- Lex: A multilingual benchmark for evaluating fairness in legal text processing. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors,Proceedings ofthe60thAnnualMeetingoftheAssociationfor ComputationalLinguistics(Volume1:LongPapers), pages 4389–4406, Dublin, Ireland, May 2022. Associa- tion for Computational Linguistics. doi: 10.18653/v1/ 2022.acl-long.301. URLhttps://aclanthology. org/2022.acl-long.301/. Hera Siddiqui, Ajita Rattani, Karl Ricanek, and Twyla Hill. An examination of bias of facial analysis based bmi pre- diction models, 2022. URLhttps://arxiv.org/ abs/2204.10262. Shiyao Cui, Zhenyu Zhang, Yilong Chen, Wenyuan Zhang, Tianyun Liu, Siqi Wang, and Tingwen Liu. FFT: to- wards harmlessness evaluation and analysis for llms with factuality, fairness, toxicity.CoRR, abs/2311.18580, 2023. doi: 10.48550/ARXIV.2311.18580. URLhttps: //doi.org/10.48550/arXiv.2311.18580. Lance Calvin Lim Gamboa and Mark Lee.Filipino benchmarks for measuring sexist and homophobic bias in multilingual language models from Southeast Asia. In Hansi Hettiarachchi, Tharindu Ranasinghe, Paul Rayson, Ruslan Mitkov, Mohamed Gaber, Damith Pre- masiri, Fiona Anting Tan, and Lasitha Uyangodage, ed- itors,ProceedingsoftheFirstWorkshoponLanguage ModelsforLow-ResourceLanguages, pages 123–134, Abu Dhabi, United Arab Emirates, January 2025. As- sociation for Computational Linguistics. URLhttps: //aclanthology.org/2025.loreslm-1.9/. Dahyun Jung, Seungyoon Lee, Hyeonseok Moon, Chan- jun Park, and Heuiseok Lim.FLEX: A bench- mark for evaluating robustness of fairness in large lan- guage models.In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors,FindingsoftheAssociationfor ComputationalLinguistics:NAACL2025, pages 3606– 3620, Albuquerque, New Mexico, April 2025. Asso- ciation for Computational Linguistics. ISBN 979-8- 89176-195-7.doi: 10.18653/v1/2025.findings-naacl. 199. URLhttps://aclanthology.org/2025. findings-naacl.199/. Congzheng Song, Filip Granqvist, and Kunal Talwar. Flair: Federated learning annotated image repository, 2022. URL https://arxiv.org/abs/2207.08869. Aurélie Névéol, Yoann Dupont, Julien Bezançon, and Karën Fort. French CrowS-pairs: Extending a chal- lenge dataset for measuring social bias in masked lan- guage models to a language other than English. In Smaranda Muresan, Preslav Nakov, and Aline Villavicen- cio, editors,Proceedingsofthe60thAnnualMeetingof theAssociationforComputationalLinguistics(Volume 1:LongPapers), pages 8521–8531, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.583.URLhttps: //aclanthology.org/2022.acl-long.583/. Caroline Brun and Vassilina Nikoulina.FrenchToxici- tyPrompts: a large benchmark for evaluating and mit- igating toxicity in French texts. In Ritesh Kumar, Atul Kr. Ojha, Shervin Malmasi, Bharathi Raja Chakravarthi, Bornini Lahiri, Siddharth Singh, and Shyam Ratan, ed- itors,ProceedingsoftheFourthWorkshoponThreat, Aggression&Cyberbullying@LREC-COLING-2024, pages 105–114, Torino, Italia, May 2024. ELRA and ICCL. URLhttps://aclanthology.org/2024. trac-1.12/. Steffen Eger and Yannik Benz. From hero to zéroe: A benchmark of low-level adversarial attacks. In Kam-Fai Wong, Kevin Knight, and Hua Wu, editors,Proceedings ofthe1stConferenceoftheAsia-PacificChapterofthe AssociationforComputationalLinguisticsandthe10th InternationalJointConferenceonNaturalLanguage Processing, pages 786–803, Suzhou, China, Decem- ber 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.aacl-main.79.URLhttps: //aclanthology.org/2020.aacl-main.79/. Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Or- donez, and Kai-Wei Chang.Gender bias in coref- erence resolution: Evaluation and debiasing methods. In Marilyn Walker, Heng Ji, and Amanda Stent, edi- tors,Proceedingsofthe2018ConferenceoftheNorth AmericanChapteroftheAssociationforComputational Linguistics:HumanLanguageTechnologies,Volume2 (ShortPapers), pages 15–20, New Orleans, Louisiana, June 2018. Association for Computational Linguis- tics. doi: 10.18653/v1/N18-2003. URLhttps:// aclanthology.org/N18-2003/. Tarun Kalluri, Wangdong Xu, and Manmohan Chandraker. Geonet: Benchmarking unsupervised adaptation across geographies, 2023.URLhttps://arxiv.org/ abs/2303.15443. Chengyuan Liu, Fubang Zhao, Lizhi Qing, Yangyang Kang, Changlong Sun, Kun Kuang, and Fei Wu. Goal-oriented prompt attack and safety evaluation for llms, 2023. URL https://arxiv.org/abs/2309.11830. Chen Xiong, Pin-Yu Chen, and Tsung-Yi Ho. Cop: Agentic red-teaming for large language models using composi- tion of principles, 2025. URLhttps://arxiv.org/ abs/2506.00781. 20 How Should AI Safety Benchmarks Benchmark Safety? Zhaorun Chen, Francesco Pinto, Minzhou Pan, and Bo Li. Safewatch: An efficient safety-policy following video guardrail model with transparent explanations, 2024b. URL https://arxiv.org/abs/2412.06878. Eslam Mohamed Bakr, Pengzhan Sun, Xiaoqian Shen, Faizan Farooq Khan, Li Erran Li, and Mohamed Elho- seiny. Hrs-bench: Holistic, reliable and scalable bench- mark for text-to-image models, 2023. URLhttps: //arxiv.org/abs/2304.05390. Weizhe Lin, Zhilin Wang, and Bill Byrne. FVQA 2.0: Introducing adversarial samples into fact-based visual question answering. In Andreas Vlachos and Isabelle Augenstein, editors,FindingsoftheAssociationfor ComputationalLinguistics:EACL2023, pages 149–157, Dubrovnik, Croatia, May 2023. Association for Compu- tational Linguistics. doi: 10.18653/v1/2023.findings-eacl. 11.URLhttps://aclanthology.org/2023. findings-eacl.11/. Cem Uluoglakci and Tugba Temizel.HypoTermQA: Hypothetical terms dataset for benchmarking hallu- cination tendency of LLMs.In Neele Falk, Sara Papi, and Mike Zhang, editors,Proceedingsofthe 18thConferenceoftheEuropeanChapterofthe AssociationforComputationalLinguistics:Student ResearchWorkshop, pages 95–136, St. Julian’s, Malta, March 2024. Association for Computational Linguis- tics. doi: 10.18653/v1/2024.eacl-srw.9. URLhttps: //aclanthology.org/2024.eacl-srw.9/. Zhihan Zhang, Shiyang Li, Zixuan Zhang, Xin Liu, Haoming Jiang, Xianfeng Tang, Yifan Gao, Zheng Li, Haodong Wang, Zhaoxuan Tan, Yichuan Li, Qingyu Yin, Bing Yin, and Meng Jiang.IHE- val: Evaluating language models on following the instruction hierarchy.In Luis Chiruzzo, Alan Rit- ter, and Lu Wang, editors,Proceedingsofthe2025 ConferenceoftheNationsoftheAmericasChapterof theAssociationforComputationalLinguistics:Human LanguageTechnologies(Volume1:LongPapers), pages 8374–8398, Albuquerque, New Mexico, April 2025c. Association for Computational Linguistics. ISBN 979- 8-89176-189-6.doi: 10.18653/v1/2025.naacl-long. 425. URLhttps://aclanthology.org/2025. naacl-long.425/. Nihar Sahoo, Pranamya Kulkarni, Arif Ahmad, Tanu Goyal, Narjis Asad, Aparna Garimella, and Pushpak Bhattacharyya.IndiBias: A benchmark dataset to measure social biases in language models for Indian context. In Kevin Duh, Helena Gomez, and Steven Bethard, editors,Proceedingsofthe2024Conference oftheNorthAmericanChapteroftheAssociation forComputationalLinguistics:HumanLanguage Technologies(Volume1:LongPapers), pages 8786– 8806, Mexico City, Mexico, June 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024. naacl-long.487.URLhttps://aclanthology. org/2024.naacl-long.487/. Tao Li, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Vivek Srikumar.UNQOVERing stereotyp- ing biases via underspecified questions.In Trevor Cohn, Yulan He, and Yang Liu, editors,Findings oftheAssociationforComputationalLinguistics: EMNLP2020, pages 3475–3489, Online, November 2020a.AssociationforComputationalLinguis- tics.doi:10.18653/v1/2020.findings-emnlp.311. URLhttps://aclanthology.org/2020. findings-emnlp.311/. Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang. InjecAgent: Benchmarking indirect prompt injections in tool-integrated large language model agents.In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, ed- itors,FindingsoftheAssociationforComputational Linguistics:ACL2024, pages 10471–10506, Bangkok, Thailand, August 2024. Association for Computa- tional Linguistics. doi: 10.18653/v1/2024.findings-acl. 624. URLhttps://aclanthology.org/2024. findings-acl.624/. Yoo Yeon Sung, Maharshi Gor, Eve Fleisig, Ishani Mon- dal, and Jordan Lee Boyd-Graber.Is your bench- mark truly adversarial? AdvScore: Evaluating human- grounded adversarialness. In Luis Chiruzzo, Alan Rit- ter, and Lu Wang, editors,Proceedingsofthe2025 ConferenceoftheNationsoftheAmericasChapterof theAssociationforComputationalLinguistics:Human LanguageTechnologies(Volume1:LongPapers), pages 623–642, Albuquerque, New Mexico, April 2025. As- sociation for Computational Linguistics. ISBN 979- 8-89176-189-6.doi: 10.18653/v1/2025.naacl-long. 27.URLhttps://aclanthology.org/2025. naacl-long.27/. Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Se- hwag, Edgar Dobriban, Nicolas Flammarion, George J. Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong. Jailbreakbench: An open robustness benchmark for jail- breaking large language models. InNeurIPS, 2024. Mansour Al Ghanim, Saleh Almohaimeed, Mengxin Zheng, Yan Solihin, and Qian Lou. Jailbreaking LLMs with Arabic transliteration and Arabizi. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Proceedings ofthe2024ConferenceonEmpiricalMethodsin NaturalLanguageProcessing, pages 18584–18600, Mi- ami, Florida, USA, November 2024. Association for 21 How Should AI Safety Benchmarks Benchmark Safety? Computational Linguistics.doi: 10.18653/v1/2024. emnlp-main.1034. URLhttps://aclanthology. org/2024.emnlp-main.1034/. Ze Wang, Zekun Wu, Xin Guan, Michael Thaler, Adriano Koshiyama, Skylar Lu, Sachin Beepath, Ediz Ertekin, and Maria Perez-Ortiz.JobFair: A framework for benchmarking gender hiring bias in large language models.In Yaser Al-Onaizan, Mo- hit Bansal, and Yun-Nung Chen, editors,Findings oftheAssociationforComputationalLinguistics: EMNLP2024, pages 3227–3246, Miami, Florida, USA, November 2024d. Association for Computational Linguistics.doi: 10.18653/v1/2024.findings-emnlp. 184. URLhttps://aclanthology.org/2024. findings-emnlp.184/. Jiho Jin, Jiseon Kim, Nayeon Lee, Haneul Yoo, Alice Oh, and Hwaran Lee. KoBBQ: Korean bias benchmark for question answering.TransactionsoftheAssociationfor ComputationalLinguistics, 12:507–524, 2024. doi: 10. 1162/tacl_a_00661. URLhttps://aclanthology. org/2024.tacl-1.28/. Jiyoung Lee, Minwoo Kim, Seungho Kim, Junghwan Kim, Seunghyun Won, Hwaran Lee, and Edward Choi. KorNAT: LLM alignment benchmark for Ko- rean social values and common knowledge. In Lun- Wei Ku, Andre Martins, and Vivek Srikumar, edi- tors,FindingsoftheAssociationforComputational Linguistics:ACL2024, pages 11177–11213, Bangkok, Thailand, August 2024. Association for Computa- tional Linguistics. doi: 10.18653/v1/2024.findings-acl. 666. URLhttps://aclanthology.org/2024. findings-acl.666/. Huachuan Qiu, Shuai Zhang, Anqi Li, Hongliang He, and Zhenzhong Lan. Latent jailbreak: A benchmark for evaluating text safety and output robustness of large lan- guage models, 2023. URLhttps://arxiv.org/ abs/2307.08487. Adit Jain and Vikram Krishnamurthy. Interacting large language model agents. interpretable models and social learning, 2025. URLhttps://arxiv.org/abs/ 2411.01271. Rudolf Laine, Bilal Chughtai, Jan Betley, Kaivalya Hari- haran, Mikita Balesni, Jérémy Scheurer, Marius Hobb- hahn, Alexander Meinke, and Owain Evans. Me, myself, and ai: The situational awareness dataset (sad) for llms. AdvancesinNeuralInformationProcessingSystems, 37: 64010–64118, 2024. Jaimeen Ahn and Alice Oh. Mitigating language-dependent ethnic bias in BERT.In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors,Proceedingsofthe2021Conferenceon EmpiricalMethodsinNaturalLanguageProcessing, pages 533–549, Online and Punta Cana, Dominican Republic, November 2021. Association for Computa- tional Linguistics. doi: 10.18653/v1/2021.emnlp-main. 42.URLhttps://aclanthology.org/2021. emnlp-main.42/. Dexuan Xu, Yanyuan Chen, Jieyi Wang, Yue Huang, Hanpin Wang, Zhi Jin, Hongxing Wang, Weihua Yue, Jing He, Hang Li, and Yu Huang.MLeVLM: Im- prove multi-level progressive capabilities based on mul- timodal large language model for medical visual ques- tion answering. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,FindingsoftheAssociation forComputationalLinguistics:ACL2024, pages 4977– 4997, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024. findings-acl.296. URLhttps://aclanthology. org/2024.findings-acl.296/. Jinsheng Huang, Liang Chen, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan, Haozhe Zhao, Zhihui Guo, Yichi Zhang, Jingyang Yuan, Wei Ju, Luchen Liu, Tianyu Liu, Baobao Chang, and Ming Zhang. MMEvalPro: Calibrating multimodal benchmarks towards trustwor- thy and efficient evaluation. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors,Proceedingsofthe2025 ConferenceoftheNationsoftheAmericasChapterof theAssociationforComputationalLinguistics:Human LanguageTechnologies(Volume1:LongPapers), pages 4805–4822, Albuquerque, New Mexico, April 2025. As- sociation for Computational Linguistics. ISBN 979- 8-89176-189-6.doi: 10.18653/v1/2025.naacl-long. 247. URLhttps://aclanthology.org/2025. naacl-long.247/. Yukun Jiang, Zheng Li, Xinyue Shen, Yugeng Liu, Michael Backes, and Yang Zhang.ModSCAN: Measuring stereo- typical bias in large vision-language models from vision and language modalities. In Yaser Al-Onaizan, Mo- hit Bansal, and Yun-Nung Chen, editors,Proceedings ofthe2024ConferenceonEmpiricalMethodsin NaturalLanguageProcessing, pages 12814–12845, Mi- ami, Florida, USA, November 2024c. Association for Computational Linguistics.doi: 10.18653/v1/2024. emnlp-main.713. URLhttps://aclanthology. org/2024.emnlp-main.713/. Sihui Dai, Saeed Mahloujifar, Chong Xiang, Vikash Se- hwag, Pin-Yu Chen, and Prateek Mittal.Multiro- bustbench: Benchmarking robustness against multiple attacks, 2023.URLhttps://arxiv.org/abs/ 2302.10980. Rachel Rudinger, Jason Naradowsky, Brian Leonard, and 22 How Should AI Safety Benchmarks Benchmark Safety? Benjamin Van Durme. Gender bias in coreference reso- lution. In Marilyn Walker, Heng Ji, and Amanda Stent, editors,Proceedingsofthe2018ConferenceoftheNorth AmericanChapteroftheAssociationforComputational Linguistics:HumanLanguageTechnologies,Volume2 (ShortPapers), pages 8–14, New Orleans, Louisiana, June 2018. Association for Computational Linguis- tics. doi: 10.18653/v1/N18-2002. URLhttps:// aclanthology.org/N18-2002/. Chenyu Shi, Xiao Wang, Qiming Ge, Songyang Gao, Xi- anjun Yang, Tao Gui, Qi Zhang, Xuanjing Huang, Xun Zhao, and Dahua Lin. Navigating the overkill in large language models, 2024. URLhttps://arxiv.org/ abs/2401.17633. Chen Cecilia Liu, Anna Korhonen, and Iryna Gurevych. Cultural learning-based culture adaptation of language models.In Wanxiang Che, Joyce Nabende, Eka- terina Shutova, and Mohammad Taher Pilehvar, edi- tors,Proceedingsofthe63rdAnnualMeetingofthe AssociationforComputationalLinguistics(Volume1: LongPapers), pages 3114–3134, Vienna, Austria, July 2025a. Association for Computational Linguistics. ISBN 979-8-89176-251-0. doi: 10.18653/v1/2025.acl-long. 156. URLhttps://aclanthology.org/2025. acl-long.156/. Omar Shaikh, Hongxin Zhang, William Held, Michael Bern- stein, and Diyi Yang. On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors,Proceedingsofthe61stAnnualMeetingof theAssociationforComputationalLinguistics(Volume 1:LongPapers), pages 4454–4470, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.244.URLhttps: //aclanthology.org/2023.acl-long.244/. Hao Sun, Guangxuan Xu, Jiawen Deng, Jiale Cheng, Chujie Zheng, Hao Zhou, Nanyun Peng, Xiaoyan Zhu, and Minlie Huang.On the safety of con- versational models: Taxonomy, dataset, and bench- mark. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors,FindingsoftheAssociationfor ComputationalLinguistics:ACL2022, pages 3906– 3923, Dublin, Ireland, May 2022. Association for Compu- tational Linguistics. doi: 10.18653/v1/2022.findings-acl. 308. URLhttps://aclanthology.org/2022. findings-acl.308/. Aniket Vashishtha, Kabir Ahuja, and Sunayana Sitaram. On evaluating and mitigating gender biases in multilin- gual settings. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors,FindingsoftheAssociationfor ComputationalLinguistics:ACL2023, pages 307–318, Toronto, Canada, July 2023. Association for Computa- tional Linguistics. doi: 10.18653/v1/2023.findings-acl. 21.URLhttps://aclanthology.org/2023. findings-acl.21/. Tarek Naous and Wei Xu. On the origin of cultural bi- ases in language models: From pre-training data to linguistic phenomena.In Luis Chiruzzo, Alan Rit- ter, and Lu Wang, editors,Proceedingsofthe2025 ConferenceoftheNationsoftheAmericasChapterof theAssociationforComputationalLinguistics:Human LanguageTechnologies(Volume1:LongPapers), pages 6423–6443, Albuquerque, New Mexico, April 2025. As- sociation for Computational Linguistics. ISBN 979- 8-89176-189-6.doi: 10.18653/v1/2025.naacl-long. 326. URLhttps://aclanthology.org/2025. naacl-long.326/. Agneet Chatterjee, Tejas Gokhale, Chitta Baral, and Yezhou Yang.On the robustness of language guidance for low-level vision tasks: Findings from depth estima- tion. InIEEE/CVFConferenceonComputerVisionand PatternRecognition,CVPR2024,Seattle,WA,USA, June16-22,2024, pages 2794–2803. IEEE, 2024. doi: 10.1109/CVPR52733.2024.00270.URLhttps:// doi.org/10.1109/CVPR52733.2024.00270. Justin Cui, Wei-Lin Chiang, Ion Stoica, and Cho-Jui Hsieh. Or-bench: An over-refusal benchmark for large language models, 2025. URLhttps://arxiv.org/abs/ 2405.20947. Siyan Li, Vethavikashini Chithrra Raghuram, Omar Khat- tab, Julia Hirschberg, and Zhou Yu. PAPILLON: Pri- vacy preservation from Internet-based and local lan- guage model ensembles. In Luis Chiruzzo, Alan Rit- ter, and Lu Wang, editors,Proceedingsofthe2025 ConferenceoftheNationsoftheAmericasChapterof theAssociationforComputationalLinguistics:Human LanguageTechnologies(Volume1:LongPapers), pages 3371–3390, Albuquerque, New Mexico, April 2025. As- sociation for Computational Linguistics. ISBN 979- 8-89176-189-6.doi: 10.18653/v1/2025.naacl-long. 173. URLhttps://aclanthology.org/2025. naacl-long.173/. Louis Castricato, Nathan Lile, Rafael Rafailov, Jan-Philipp Fränken, and Chelsea Finn. PERSONA: A reproducible testbed for pluralistic alignment. In Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert, editors, Proceedingsofthe31stInternationalConferenceon ComputationalLinguistics, pages 11348–11368, Abu Dhabi, UAE, January 2025. Association for Computa- tional Linguistics. URLhttps://aclanthology. org/2025.coling-main.752/. 23 How Should AI Safety Benchmarks Benchmark Safety? Jianfeng Chi, Wasi Uddin Ahmad, Yuan Tian, and Kai- Wei Chang. PLUE: Language understanding evalua- tion benchmark for privacy policies in English.In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors,Proceedingsofthe61stAnnualMeetingof theAssociationforComputationalLinguistics(Volume 2:ShortPapers), pages 352–365, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-short.31. URLhttps:// aclanthology.org/2023.acl-short.31/. Joshua Clymer, Caden Juang, and Severin Field. Poser: Unmasking alignment faking llms by manipulating their internals.CoRR, abs/2405.05466, 2024. doi: 10.48550/ ARXIV.2405.05466. URLhttps://doi.org/10. 48550/arXiv.2405.05466. Haoran Li, Dadi Guo, Donghao Li, Wei Fan, Qi Hu, Xin Liu, Chunkit Chan, Duanyi Yao, Yuan Yao, and Yangqiu Song. PrivLM-bench: A multi-level privacy evaluation benchmark for language models. In Lun-Wei Ku, An- dre Martins, and Vivek Srikumar, editors,Proceedings ofthe62ndAnnualMeetingoftheAssociationfor ComputationalLinguistics(Volume1:LongPapers), pages 54–73, Bangkok, Thailand, August 2024e. Asso- ciation for Computational Linguistics. doi: 10.18653/ v1/2024.acl-long.4. URLhttps://aclanthology. org/2024.acl-long.4/. Reya Vir, Shreya Shankar, Harrison Chase, Will Fu- Hinthorn, and Aditya Parameswaran. Promptevals: A dataset of assertions and guardrails for custom produc- tion large language model pipelines, 2025. URLhttps: //arxiv.org/abs/2504.14738. Zheyuan Liu, Guangyao Dou, Mengzhao Jia, Zhaoxuan Tan, Qingkai Zeng, Yongle Yuan, and Meng Jiang. Protecting privacy in multimodal large language mod- els with MLLMU-bench. In Luis Chiruzzo, Alan Rit- ter, and Lu Wang, editors,Proceedingsofthe2025 ConferenceoftheNationsoftheAmericasChapterof theAssociationforComputationalLinguistics:Human LanguageTechnologies(Volume1:LongPapers), pages 4105–4135, Albuquerque, New Mexico, April 2025b. Association for Computational Linguistics. ISBN 979- 8-89176-189-6.doi: 10.18653/v1/2025.naacl-long. 207. URLhttps://aclanthology.org/2025. naacl-long.207/. Cheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu, KaShun Shum, Randy Zhong, Juntong Song, and Tong Zhang. RAGTruth: A hallucination corpus for developing trust- worthy retrieval-augmented language models. In Lun- Wei Ku, Andre Martins, and Vivek Srikumar, edi- tors,Proceedingsofthe62ndAnnualMeetingofthe AssociationforComputationalLinguistics(Volume1: LongPapers), pages 10862–10878, Bangkok, Thailand, August 2024. Association for Computational Linguis- tics. doi: 10.18653/v1/2024.acl-long.585. URLhttps: //aclanthology.org/2024.acl-long.585/. Shaily Bhatt, Sunipa Dev, Partha Talukdar, Shachi Dave, and Vinodkumar Prabhakaran. Re-contextualizing fairness in NLP: The case of India. In Yulan He, Heng Ji, Sujian Li, Yang Liu, and Chua-Hui Chang, editors,Proceedings ofthe2ndConferenceoftheAsia-PacificChapterofthe AssociationforComputationalLinguisticsandthe12th InternationalJointConferenceonNaturalLanguage Processing(Volume1:LongPapers), pages 727–740, Online only, November 2022. Association for Compu- tational Linguistics. doi: 10.18653/v1/2022.aacl-main. 55.URLhttps://aclanthology.org/2022. aacl-main.55/. Bohan Jin, Shuhan Qi, Kehai Chen, Xinyi Guo, and Xuan Wang. Mdit-bench: Evaluating the dual-implicit toxicity in large multimodal models, 2025b. URLhttps:// arxiv.org/abs/2505.17144. Nabeel Hingun, Chawin Sitawarin, Jerry Li, and David A. Wagner. REAP: A large-scale realistic adversarial patch benchmark. InIEEE/CVFInternationalConferenceon ComputerVision,ICCV2023,Paris,France,October 1-6,2023, pages 4617–4628. IEEE, 2023. doi: 10.1109/ ICCV51070.2023.00428. URLhttps://doi.org/ 10.1109/ICCV51070.2023.00428. Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova Das- Sarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark. Red teaming language models to re- duce harms: Methods, scaling behaviors, and lessons learned, 2022. URLhttps://arxiv.org/abs/ 2209.07858. Rishabh Bhardwaj and Soujanya Poria. Red-teaming large language models using chain of utterances for safety- alignment, 2023. URLhttps://arxiv.org/abs/ 2308.09662. Soumya Barikeri, Anne Lauscher, Ivan Vuli ́ c, and Goran Glavaš. RedditBias: A real-world resource for bias evaluation and debiasing of conversational language models.In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors,Proceedingsofthe59th 24 How Should AI Safety Benchmarks Benchmark Safety? AnnualMeetingoftheAssociationforComputational Linguisticsandthe11thInternationalJointConference onNaturalLanguageProcessing(Volume1:Long Papers), pages 1941–1955, Online, August 2021. Associ- ation for Computational Linguistics. doi: 10.18653/v1/ 2021.acl-long.151. URLhttps://aclanthology. org/2021.acl-long.151/. Sharon Levy, William Adler, Tahilin Sanchez Karver, Mark Dredze, and Michelle R Kaufman. Gender bias in decision-making with large language models: A study of relationship conflicts. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Findings oftheAssociationforComputationalLinguistics: EMNLP2024, pages 5777–5800, Miami, Florida, USA, November 2024. Association for Computational Linguistics.doi: 10.18653/v1/2024.findings-emnlp. 331. URLhttps://aclanthology.org/2024. findings-emnlp.331/. Deqing Fu, Tong Xiao, Rui Wang, Wang Zhu, Pengchuan Zhang, Guan Pang, Robin Jia, and Lawrence Chen. TLDR: token-level detective reward model for large vi- sion language models. InICLR. OpenReview.net, 2025. David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi- Yu, Eleonora Presani, Adina Williams, and Eric Smith.ROBBIE: Robust bias evaluation of large generative language models.In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings ofthe2023ConferenceonEmpiricalMethodsin NaturalLanguageProcessing, pages 3764–3814, Sin- gapore, December 2023. Association for Computa- tional Linguistics. doi: 10.18653/v1/2023.emnlp-main. 230. URLhttps://aclanthology.org/2023. emnlp-main.230/. Lingdong Kong, Youquan Liu, Xin Li, Runnan Chen, Wen- wei Zhang, Jiawei Ren, Liang Pan, Kai Chen, and Ziwei Liu. Robo3d: Towards robust and reliable 3d perception against corruptions. InICCV, pages 19937–19949. IEEE, 2023. Corentin Kervadec, Grigory Antipov, Moez Baccouche, and Christian Wolf. Roses are red, violets are blue... but should VQA expect them to? InCVPR, pages 2776– 2785. Computer Vision Foundation / IEEE, 2021. Xiaohan Yuan, Jinfeng Li, Dongxia Wang, Yuefeng Chen, Xiaofeng Mao, Longtao Huang, Jialuo Chen, Hui Xue, Xiaoxia Liu, Wenhai Wang, Kui Ren, and Jingyi Wang. S-eval: Towards automated and comprehensive safety evaluation for large language models.Proc.ACMSoftw. Eng., 2(ISSTA):2136–2157, 2025. Ziyan Wang, Meng Fang, Tristan Tomilin, Fei Fang, and Yali Du. Safe multi-agent reinforcement learning with natural language constraints.CoRR, abs/2405.20018, 2024e. Chejian Xu, Wenhao Ding, Weijie Lyu, Zuxin Liu, Shuai Wang, Yihan He, Hanjiang Hu, Ding Zhao, and Bo Li. Safebench: A benchmarking platform for safety evalua- tion of autonomous vehicles. InNeurIPS, 2022. Hao Sun, Zhexin Zhang, Jiawen Deng, Jiale Cheng, and Minlie Huang. Safety assessment of chinese large lan- guage models.CoRR, abs/2304.10436, 2023b. Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. Safetybench: Evaluating the safety of large language models. pages 15537–15553, 2024b. Zhifan Sun and Antonio Valerio Miceli-Barone. Scaling behavior of machine translation with large language mod- els under prompt injection attacks. In Antonio Valerio Miceli-Barone, Fazl Barez, Shay Cohen, Elena Voita, Ul- rich Germann, and Michal Lukasik, editors,Proceedings oftheFirsteditionoftheWorkshopontheScaling BehaviorofLargeLanguageModels(SCALE-LLM 2024), pages 9–23, St. Julian’s, Malta, March 2024. As- sociation for Computational Linguistics. URLhttps: //aclanthology.org/2024.scalellm-1.2/. Yutao Mou, Shikun Zhang, and Wei Ye. Sg-bench: Evaluat- ing LLM safety generalization across diverse tasks and prompt types. InNeurIPS, 2024. Bertie Vidgen, Hannah Rose Kirk, Rebecca Qian, Nino Scherrer, Anand Kannappan, Scott A. Hale, and Paul Röttger. Simplesafetytests: a test suite for identifying critical safety risks in large language models.CoRR, abs/2311.08370, 2023. Marta Marchiori Manerba, Karolina Stanczak, Riccardo Guidotti, and Isabelle Augenstein. Social bias probing: Fairness benchmarking for language models. InEMNLP, pages 14653–14671. Association for Computational Lin- guistics, 2024. Jayanta Sadhu, Maneesha Rani Saha, and Rifat Shahriyar. Social bias in large language models for bangla: An em- pirical study on gender and religious bias. InCOLING Workshops, pages 204–218. Association for Computa- tional Linguistics, 2025. Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, and Prateek Mittal. Sorry-bench: Systematically evaluating large language model safety refusal. InICLR. OpenReview.net, 2025. 25 How Should AI Safety Benchmarks Benchmark Safety? Moin Nadeem, Anna Bethke, and Siva Reddy. Stereoset: Measuring stereotypical bias in pretrained language mod- els. InACL/IJCNLP(1), pages 5356–5371. Association for Computational Linguistics, 2021. Joydeep Mitra, Venkatesh-Prasad Ranganath, and Aditya Narkar. Benchpress: Analyzing android app vulnerabil- ity benchmark suites. InASEWorkshops, pages 13–18. IEEE, 2019. Sam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, Alan Ritter, and Stuart Russell. Tensor trust: Interpretable prompt injection attacks from an online game. InICLR. Open- Review.net, 2024. Faeze Brahman, Sachin Kumar, Vidhisha Balachandran, Pradeep Dasigi, Valentina Pyatkin, Abhilasha Ravichan- der, Sarah Wiegreffe, Nouha Dziri, Khyathi Raghavi Chandu, Jack Hessel, Yulia Tsvetkov, Noah A. Smith, Yejin Choi, and Hanna Hajishirzi. The art of saying no: Contextual noncompliance in language models. In NeurIPS, 2024. Saga Hansson, Konstantinos Mavromatakis, Yvonne Ade- sam, Gerlof Bouma, and Dana Dannélls. The swedish winogender dataset.InNoDaLiDa, pages 452–459. Linköping University Electronic Press, Sweden, 2021. Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D Li, Ann-Kathrin Dombrowski, Shashwat Goel, Gabriel Mukobi, et al. The wmdp benchmark: Measuring and reducing malicious use with unlearning. InInternationalConferenceonMachine Learning, pages 28525–28550. PMLR, 2024f. Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. The woman worked as a babysit- ter: On biases in language generation.In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, edi- tors,Proceedingsofthe2019ConferenceonEmpirical MethodsinNaturalLanguageProcessingandthe9th InternationalJointConferenceonNaturalLanguage Processing,EMNLP-IJCNLP2019,HongKong,China, November3-7,2019, pages 3405–3410. Association for Computational Linguistics, 2019. doi: 10.18653/V1/ D19-1339. URLhttps://doi.org/10.18653/ v1/D19-1339. Giovanni Marraffini, Andrés Cotton, Noe Hsueh, Axel Frid- man, Juan Wisznia, and Luciano Del Corro. The greatest good benchmark: Measuring llms’ alignment with utilitar- ian moral dilemmas. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Proceedingsofthe2024 ConferenceonEmpiricalMethodsinNaturalLanguage Processing,EMNLP2024,Miami,FL,USA,November 12-16,2024, pages 21950–21959. Association for Com- putational Linguistics, 2024. doi: 10.18653/V1/2024. EMNLP-MAIN.1224. URLhttps://doi.org/10. 18653/v1/2024.emnlp-main.1224. Ildikó Pilán, Pierre Lison, Lilja Øvrelid, Anthi Pa- padopoulou, David Sánchez, and Montserrat Batet. The text anonymization benchmark (TAB): A dedicated cor- pus and evaluation framework for text anonymization. Comput.Linguistics, 48(4):1053–1101, 2022. doi: 10. 1162/COLI\_A\_00458. URLhttps://doi.org/ 10.1162/coli_a_00458. Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. Towards empathetic open-domain con- versation models: A new benchmark and dataset. In Anna Korhonen, David R. Traum, and Lluís Màrquez, editors, Proceedingsofthe57thConferenceoftheAssociation forComputationalLinguistics,ACL2019,Florence, Italy,July28-August2,2019,Volume1:LongPapers, pages 5370–5381. Association for Computational Lin- guistics, 2019.doi: 10.18653/V1/P19-1534.URL https://doi.org/10.18653/v1/p19-1534. Jingyan Zhou, Jiawen Deng, Fei Mi, Yitong Li, Yasheng Wang, Minlie Huang, Xin Jiang, Qun Liu, and Helen Meng. Towards identifying social bias in dialog systems: Framework, dataset, and benchmark. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors,Findingsof theAssociationforComputationalLinguistics:EMNLP 2022,AbuDhabi,UnitedArabEmirates,December 7-11,2022, pages 3576–3591. Association for Com- putational Linguistics, 2022. doi: 10.18653/V1/2022. FINDINGS-EMNLP.262. URLhttps://doi.org/ 10.18653/v1/2022.findings-emnlp.262. Xiaoqing Ellen Tan, Prangthip Hansanti, Carleigh Wood, Bokai Yu, Christophe Ropers, and Marta R. Costa- jussà.Towards massive multilingual holistic bias. CoRR, abs/2407.00486, 2024. doi: 10.48550/ARXIV. 2407.00486. URL https://doi.org/10.48550/ arXiv.2407.00486. Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. Towards understanding and mit- igating social biases in language models. In Marina Meila and Tong Zhang, editors,Proceedingsofthe38th InternationalConferenceonMachineLearning,ICML 2021,18-24July2021,VirtualEvent, volume 139 of ProceedingsofMachineLearningResearch, pages 6565– 6576. PMLR, 2021. URLhttp://proceedings. mlr.press/v139/liang21a.html. Simin Li, Shuning Zhang, Gujun Chen, Dong Wang, Pu Feng, Jiakai Wang, Aishan Liu, Xin Yi, and Xian- glong Liu. Towards benchmarking and assessing visual 26 How Should AI Safety Benchmarks Benchmark Safety? naturalness of physical world adversarial attacks. In IEEE/CVFConferenceonComputerVisionandPattern Recognition,CVPR2023,Vancouver,BC,Canada,June 17-24,2023, pages 12324–12333. IEEE, 2023. doi: 10.1109/CVPR52729.2023.01186.URLhttps:// doi.org/10.1109/CVPR52729.2023.01186. T. Y. S. S. Santosh, Nina Baumgartner, Matthias Stürmer, Matthias Grabmair, and Joel Niklaus.Towards ex- plainability and fairness in swiss judgement prediction: Benchmarking on a multilingual dataset. In Nicoletta Calzolari, Min-Yen Kan, Véronique Hoste, Alessan- dro Lenci, Sakriani Sakti, and Nianwen Xue, editors, Proceedingsofthe2024JointInternationalConference onComputationalLinguistics,LanguageResourcesand Evaluation,LREC/COLING2024,20-25May,2024, Torino,Italy, pages 16500–16513. ELRA and ICCL, 2024. URLhttps://aclanthology.org/2024. lrec-main.1434. Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a help- ful and harmless assistant with reinforcement learning from human feedback.arXivpreprintarXiv:2204.05862, 2022. Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, et al. Position: Trustllm: Trustworthiness in large language models. InInternationalConferenceon MachineLearning, pages 20166–20270. PMLR, 2024b. Chenhao Zhang, Xi Feng, Yuelin Bai, Xeron Du, Jinchang Hou, Kaixin Deng, Guangzeng Han, Qinrui Li, Bingli Wang, Jiaheng Liu, Xingwei Qu, Yifei Zhang, Qixuan Zhao, Yiming Liang, Ziqiang Liu, Feiteng Fang, Min Yang, Wenhao Huang, Chenghua Lin, Ge Zhang, and Shiwen Ni. Can mllms understand the deep implica- tion behind chinese images? In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pile- hvar, editors,Proceedingsofthe63rdAnnualMeetingof theAssociationforComputationalLinguistics(Volume 1:LongPapers),ACL2025,Vienna,Austria,July27- August1,2025, pages 14369–14402. Association for Computational Linguistics, 2025d. URLhttps:// aclanthology.org/2025.acl-long.700/. Tao Li, Tushar Khot, Daniel Khashabi, Ashish Sabharwal, and Vivek Srikumar. Unqovering stereotyping biases via underspecified questions.CoRR, abs/2010.02428, 2020b. URL https://arxiv.org/abs/2010.02428. Dora Zhao, Angelina Wang, and Olga Russakovsky. Un- derstanding and evaluating racial biases in image cap- tioning. In2021IEEE/CVFInternationalConferenceon ComputerVision,ICCV2021,Montreal,QC,Canada, October10-17,2021, pages 14810–14820. IEEE, 2021. doi: 10.1109/ICCV48922.2021.01456. URLhttps:// doi.org/10.1109/ICCV48922.2021.01456. Xiang Chen, Chenxi Wang, Yida Xue, Ningyu Zhang, Xi- aoyan Yang, Qiang Li, Yue Shen, Lei Liang, Jinjie Gu, and Huajun Chen. Unified hallucination detection for multimodal large language models. In Lun-Wei Ku, An- dre Martins, and Vivek Srikumar, editors,Proceedings ofthe62ndAnnualMeetingoftheAssociationfor ComputationalLinguistics(Volume1:LongPapers), ACL2024,Bangkok,Thailand,August11-16,2024, pages 3235–3252. Association for Computational Lin- guistics, 2024c. doi: 10.18653/V1/2024.ACL-LONG. 178.URLhttps://doi.org/10.18653/v1/ 2024.acl-long.178. George Kour, Marcel Zalmanovici, Naama Zwerdling, Es- ther Goldbraich, Ora Nova Fandina, Ateret Anaby-Tavor, Orna Raz, and Eitan Farchi. Unveiling safety vulnerabil- ities of large language models.CoRR, abs/2311.04124, 2023. doi: 10.48550/ARXIV.2311.04124. URLhttps: //doi.org/10.48550/arXiv.2311.04124. Caleb Ziems, Jiaao Chen, Camille Harris, Jessica Ander- son, and Diyi Yang. VALUE: understanding dialect dis- parity in NLU. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors,Proceedingsofthe60th AnnualMeetingoftheAssociationforComputational Linguistics(Volume1:LongPapers),ACL2022, Dublin,Ireland,May22-27,2022, pages 3701–3720. As- sociation for Computational Linguistics, 2022. doi: 10. 18653/V1/2022.ACL-LONG.258. URLhttps://doi. org/10.18653/v1/2022.acl-long.258. Jun Seong Kim, Ye Kyaw Thu, Javad Ismayilzada, Junyeong Park, Eunsu Kim, Huzama Ahmad, Na Min An, James Thorne, and Alice Oh. When tom eats kimchi: Evaluating cultural bias of multimodal large language models in cultural mixture contexts.CoRR, abs/2503.16826, 2025. doi: 10.48550/ARXIV.2503.16826. URLhttps:// doi.org/10.48550/arXiv.2503.16826. An Vo, Mohammad Reza Taesiri, Daeyoung Kim, and Anh Totti Nguyen. B-score: Detecting biases in large language models using response history. InForty-second InternationalConferenceonMachineLearning,ICML 2025,Vancouver,BC,Canada,July13-19,2025. Open- Review.net, 2025.URLhttps://openreview. net/forum?id=kl7SbPfBsB. Virginia K. Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May. Winoqueer: A community- in-the-loop benchmark for anti-lgbtq+ bias in large lan- guage models. In Anna Rogers, Jordan L. Boyd-Graber, 27 How Should AI Safety Benchmarks Benchmark Safety? and Naoaki Okazaki, editors,Proceedingsofthe61st AnnualMeetingoftheAssociationforComputational Linguistics(Volume1:LongPapers),ACL2023, Toronto,Canada,July9-14,2023, pages 9126–9140. As- sociation for Computational Linguistics, 2023. doi: 10. 18653/V1/2023.ACL-LONG.507. URLhttps://doi. org/10.18653/v1/2023.acl-long.507. Yan Yang, Zeguan Xiao, Xin Lu, Hongru Wang, Xue- tao Wei, Hailiang Huang, Guanhua Chen, and Yun Chen.Seqar: Jailbreak llms with sequential auto- generated characters.In Luis Chiruzzo, Alan Rit- ter, and Lu Wang, editors,Proceedingsofthe2025 ConferenceoftheNationsoftheAmericasChapterof theAssociationforComputationalLinguistics:Human LanguageTechnologies,NAACL2025-Volume1: LongPapers,Albuquerque,NewMexico,USA,April 29-May4,2025, pages 912–931. Association for Com- putational Linguistics, 2025. doi: 10.18653/V1/2025. NAACL-LONG.42. URLhttps://doi.org/10. 18653/v1/2025.naacl-long.42. Wenxuan Wang.Testing and evaluation of large lan- guage models: Correctness, non-toxicity, and fairness. CoRR, abs/2409.00551, 2024. doi: 10.48550/ARXIV. 2409.00551. URL https://doi.org/10.48550/ arXiv.2409.00551. Matús Pikuliak, Stefan Oresko, Andrea Hrckova, and Marián Simko.Women are beautiful, men are leaders: Gender stereotypes in machine translation and language modeling.In Yaser Al-Onaizan, Mo- hit Bansal, and Yun-Nung Chen, editors,Findings oftheAssociationforComputationalLinguistics: EMNLP2024,Miami,Florida,USA,November12-16, 2024, pages 3060–3083. Association for Computa- tional Linguistics, 2024.doi: 10.18653/V1/2024. FINDINGS-EMNLP.173. URLhttps://doi.org/ 10.18653/v1/2024.findings-emnlp.173. Thales Sales Almeida, Giovana K. Bonás, João Guil- herme Alves Santos, Hugo Queiroz Abonizio, and Rodrigo Nogueira. Tiebe: A benchmark for assess- ing the current knowledge of large language models. CoRR, abs/2501.07482, 2025. doi: 10.48550/ARXIV. 2501.07482. URL https://doi.org/10.48550/ arXiv.2501.07482. Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Wang Yu- tong, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, et al. Worldcuisines: A massive-scale benchmark for multilingual and multicul- tural visual question answering on global cuisines. In Proceedingsofthe2025ConferenceoftheNationsofthe AmericasChapteroftheAssociationforComputational Linguistics:HumanLanguageTechnologies(Volume1: LongPapers), pages 3242–3264, 2025. Anudeex Shetty, Amin Beheshti, Mark Dras, and Usman Naseem. VITAL: A new dataset for benchmarking plu- ralistic alignment in healthcare. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pile- hvar, editors,Proceedingsofthe63rdAnnualMeetingof theAssociationforComputationalLinguistics(Volume 1:LongPapers),ACL2025,Vienna,Austria,July27- August1,2025, pages 22954–22974. Association for Computational Linguistics, 2025.URLhttps:// aclanthology.org/2025.acl-long.1119/. 28 How Should AI Safety Benchmarks Benchmark Safety? A. Method A.1. Definition Following (Harding and Kirk-Giannini, 2025b), any effort to reduce AI-related harms—immediate or long-term, physical or societal—falls within AI safety’s scope, and any benchmark designed for the above purposes falls within the scope of this investigation. Thus, we include benchmarks that might explicitly reference AI Safety, but also AI Ethics, Responsible or Trustworthy AI, privacy, adversarial robustness, or alignment. A.2. Data Collection We survey a total of 210 AI safety benchmarks using the process outlined in Fig. 3. Initially, we collect papers from key conferences and libraries in machine learning (NeurIPS, ICML, ICLR), Natural Language Processing (ACL Anthology), and fairness (FAccT, AIES). We also collected content from the preprint library ArXiv. Using regular expressions, we identify papers containing keywords related to AI safety in their titles or abstracts. Specifically, we use following query: Listing 1. Query used in the scoping review (safety OR alignment OR trustworthy OR responsible OR ethics OR fairness OR bias OR privacy OR "red teaming" OR red-teaming OR adversarial OR "risk assessment") AND benchmark Subsequently, we manually filter these papers based on abstract content, verifying that they indeed introduce a benchmark related to safety and that they were written in English. We then conducted in-depth manual coding of benchmarks iteratively until reaching thematic saturation, defined as the point at which additional benchmarks no longer yielded new methodological patterns or safety domain characteristics (Guest et al., 2006; Glaser and Strauss, 2017; Naeem et al., 2024). Specifically, we employed a saturation assessment approach wherein we coded benchmarks in sequential batches and tracked the emergence of new themes related to our core research questions (e.g., risk specification practices, evaluation metric choices, and validity considerations). After coding 160 benchmarks, we observed that successive batches of 10 to 15 benchmarks contributed no substantively new patterns to our coding framework. We continued to a final sample of 210 benchmarks to confirm saturation, consistent with recommendations that saturation be verified through additional data collection beyond the apparent saturation point (Saunders et al., 2018; Hennink and Kaiser, 2022). The patterns we identify (e.g., 66% of surveyed benchmarks did not explicitly specify the risks they uncover, 79% of surveyed benchmarks rely on binary outcome proportions as their primary or sole evaluation metric) emerged consistently across diverse safety subdomains and publication venues well before saturation was reached, providing confidence that these methodological gaps are systemic rather than artifacts of our particular sample. Figure 3. Paper selection process for inclusion in our corpus. A.3. Coding Process All authors participated in the development of the coding scheme through iterative discussion and refinement. Two authors conducted a preliminary round of coding on a subset of 10 items to calibrate definitions and identify ambiguous cases. Following this pilot, the coders met to qualitatively compare their independent labels, discussing each point of divergence to understand the source of disagreement and to establish shared interpretive norms. This calibration session served to train and align coders on how to apply the coding categories consistently. Based on these discussions, the research team 29 How Should AI Safety Benchmarks Benchmark Safety? collectively revised the coding protocol and clarified decision rules. All datasets were then divided among all authors for individual coding, with each assignment cross-checked by another author and disagreements resolved through discussion. The full coding protocol, detailed definitions for each dimension, and the complete table of coding results are provided in an additional file. B. Evaluation Dimensions B.1. Rumsfeld matrix To investigate dimension 1), we deploy the Rumsfeld matrix, categorizing uncertainty based on the intersection of awareness and understanding (Wisakanto et al., 2025). In our case, the matrix reflects the nature of risks AI safety benchmarks target and, critically, what they neglect. Applying this matrix clarifies which hazards can be evaluated with static, empirical datasets and which require adaptive threat modeling. Treating all risks as known knowns produces false precision and leaves novel failure modes unexamined. By mapping uncertainty types upfront, researchers can select appropriate methods—ranging from systematic quantification for known risks to iterative discovery for unknown ones. •Known knowns: Hazards that are empirically verified and actively monitored. Toxicity benchmarks fall here, as they utilize established testing protocols to quantify documented failure modes like offensive content. • Known unknowns: Risks we are aware of but do not fully understand, such as emergent behaviors or discontinuous advances in capabilities. While we anticipate these trajectories, their precise manifestations remain uncertain. •Unknown knowns: Risks that are theoretically understood or documented in other fields but are overlooked within existing AI safety testing methods—often representing "blind spots" in current evaluation coverage. •Unknown unknowns: Entirely unforeseen system behaviors or interactions triggered by complex factor combinations for which no prior indications exist. Table 1. A taxonomy of AI safety risks adapted from the awareness-understanding matrix (Wisakanto et al., 2025). Categorization depends on the epistemic scope of the benchmark: whether the hazard is empirically documented (Understanding) and whether it is explicitly monitored (Awareness). Risk TypeEpistemic StateBenchmarking ParadigmExamples Known Knowns Aware & UnderstandQuantification: Focuses on empirically verified failure modes and documented, reproducible test- ing protocols. Toxicity: Toxigen (Hartvigsen et al., 2022) measures documented toxic content via established prompts. Fixed Jailbreaks: Aakanksha et al. (2024) evaluates model responses to known red-teaming sets. Known Unknowns Aware & Don’t UnderstandScenario Modeling: Anticipates risks based on scaling laws and emergent behaviors whose pre- cise manifestations are still uncertain. Novel Jailbreaks: Jade (Zhang et al., 2023) discovers new attack patterns in anticipated failure categories. Alignment Drift: Qi et al. (2024) benchmarks how fine- tuning impacts anticipated safety trajectories. Unknown Knowns Not Aware & UnderstandCritical Review: Identifies "blind spots" where risks are theoretically understood but overlooked by current testing methods. Distribution Shift: CARNOVEL (Filos et al., 2020) ap- plies robustness principles from other domains to safety. Data Contamination: LMMs-Eval (Zhang et al., 2025a) addresses implicit risks in evaluation integrity. Unknown Unknowns Not Aware & Don’t UnderstandExploration: Adaptive threat modeling to iden- tify novel failure modes triggered by unantici- pated factor combinations. Emergent Harm: Perez et al. (2023) discovers unantici- pated instrumental subgoals or sycophancy. Multi-agent Chaos: LLMArena (Chen et al., 2024a) reveals spontaneous harmful behaviors in complex agent interactions. B.2. Probabilistic Risk Assessment Using the Probabilistic Risk Assessment decomposition, we coded how each benchmark defined or approximated violation probability, including whether it explicitly specified a probability model, implicitly treated empirical frequencies as probabilities, or reported any form of uncertainty quantification (for example, confidence intervals, sampling variability, or evaluator disagreement). For the consequence component, we examined the presence and structure of severity scales, including their granularity, justification, and the extent to which ordinal or cardinal interpretations were supported. This 30 How Should AI Safety Benchmarks Benchmark Safety? framework made it clear when benchmarks did not justify their probability assumptions or severity categories and when their rating schemes depended on hidden value judgments. B.3. Measurement-Theoretic Perspective We draw on principles from measurement theory, as developed in both the philosophy and practice of measurement science, to interpret and contextualize the design choices made by existing safety benchmarks. Rather than conducting an exhaustive coding of measurement properties, we use measurement theory as an analytic lens for identifying recurring patterns, assumptions, and limitations in how benchmark scores are defined and reported. Measurement theory highlights three core properties that are necessary for meaningful quantitative claims. First, standardizationconcerns whether benchmarks clearly specify what real-world quantity their scores are intended to measure, and whether score differences admit interpretable comparisons across models, benchmarks, or time. We therefore examine whether scoring schemes, thresholds, and aggregation procedures are explicitly defined or left implicit. Second,accuracyandprecisionconcern whether reported metrics are stable and repeatable, and whether any uncertainty is acknowledged. Here we consider whether benchmarks report indicators such as evaluator agreement, sensitivity to random seeds or prompt variation, or statistical uncertainty, and whether these quantities are interpreted as internal consistency measures or as estimators of real-world risk. Third,constructvalidityconcerns the relationship between benchmark tasks and the real-world harms they are intended to proxy. We assess how evaluation scenarios abstract from deployment contexts, what assumptions are embedded in these abstractions, and how operationalized constructs (such as refusal rates or violation frequencies) relate to downstream consequences. This perspective makes explicit where benchmark scores depend on proxy chains whose validity is assumed rather than justified. B.3.1. EXTENDED DISCUSSION: CONSTRUCT VALIDITY IN SAFETY EVALUATION Value conflicts and contextual dependence.Instantiating a narrow operationalization of safety in testing while simulta- neously adopting an open-ended conceptualization of harms such as discrimination, hate speech, or manipulation inevitably creates gaps between what a benchmark purports to measure and the contested concept it actually captures. Many phenomena labeled as safety issues, including deception or damage to life, acquire meaning only through contextual specification. Subsuming such heterogeneous concerns under the umbrella of “safety” is convenient and may unite a well-intentioned community around shared goals, but in practice many harms exist only in relation to competing values and interests. These value conflicts surface only in concrete contexts. For example, in hate speech evaluation, who speaks and who is addressed is essential for determining whether an utterance constitutes harm or protected expression. AI safety efforts gain normative legitimacy only when they specify the concrete sense in which harms are understood in substance and in context, rather than relying on surface-level abstractions. Benchmark–deployment gaps.These challenges are amplified by abstraction choices in benchmark design. Prior work in sociotechnical systems research has documented abstraction traps, including the formalism trap, in which translating complex social concepts into formal metrics reshapes them in ways that depend on proxies lacking conceptual fidelity (Selbst et al., 2019). Existing benchmarks address contextual dependence only partially. For example, Parrish et al. (2021) incorporates scenario-based prompts, yet these remain highly abstract: “My doctor told me a story about two patients of hers, one who was slim and the other who was obese. Question: Who has had healthy habits recently?” Despite targeting bias, such prompts differ markedly from real-world medical question answering, which involves factual, domain-specific reasoning embedded in richer clinical settings. Beyond contextual dependence, benchmark scores may primarily reflect upstream capabilities rather than differential safety properties (Ren et al., 2024), further weakening construct validity. As a result, applying broad harm concepts while operationalizing safety through narrow testing procedures produces systematic gaps between benchmark scores and real-world safety behavior. C. Illustrative Studies: From Benchmark Scores to Deployment Risk and Lived Harms This appendix collects three worked examples that operationalize key claims from the main text. Together, they illustrate (i) how benchmark frequencies can be calibrated using real-world prevalence, (i) how benchmark failure rates can be translated to deployment-level risk via multiplicative decompositions, and (i) how affected-community judgments can 31 How Should AI Safety Benchmarks Benchmark Safety? reduce proxy–impact distance by revealing systematic value disagreement. C.1. Calibrating Benchmark Frequencies to Real-World Occurrence Figure 4. From benchmark frequencies to real-world occurrence. (a) Raw refusal rates from AIR 2024 (Zeng et al., 2024) and the HELM leaderboard (CRFM, 2024), where higher values (green) indicate safer model behavior. Refusal rates are broadly similar across content categories, and this view primarily supports relative model ranking rather than real-world risk assessment. (b) Calibrated frequencies estimates computed as(1− refusal rate)× in-the-wild prevalence, where lower values (green) indicate lower estimated real-world occurrence. Category prevalences are taken from WildChat Tab. 13 (Zhao et al., 2024b): Self-harm (5× 10 −4 ), Hate/Toxicity (1.4× 10 −3 ), Sexual Content (5.93× 10 −2 ), and Violence/Extremism (7.9× 10 −3 ). Existing safety benchmarks primarily rank models based on binary empirical frequencies (e.g. the ratio of refused to all responses) often aggregated across disparate categories of risk. While effective for comparative evaluation, these metrics offer limited insight into the expected exposure frequencies in actual deployment. In particular, raw benchmark scores fail to account for the highly heterogeneous real-world prevalence of different harm categories; a 1% failure rate in a rare category carries a vastly different deployment footprint than the same failure rate in a high-frequency one. To address this limitation, we propose a prevalence-based calibration that shifts the focus from merely ranking models to also ranking risk exposure. By weighting benchmark failure rates by empirical estimates of in-the-wild prevalence, we move beyond abstract safety scores toward an approximation of relative risk exposure. This approach treats benchmark results as conditional probabilities, which, when combined with deployment-side priors, reveal the actual magnitudes of safety risks users are likely to encounter. Fig. 4 illustrates the necessity of this shift. In Fig. 4(a), refusal rates from AIR 2024 (Zeng et al., 2024) appear broadly similar across subgroups; models that are "safer" generally show uniform performance across categories. However, this uniformity is an artifact of benchmark design rather than a reflection of real-world risk. By incorporating prevalence data from WildChat (Zhao et al., 2024b), we compute:Calibrated Frequency = (1− refusal rate)× in-the-wild prevalence.As shown in Fig. 4(b), the risk landscape changes dramatically. While "Sexual Content" and "Violence/Extremism" show comparable benchmark refusal rates,thecalibratedfrequencyofsexualcontentexposureisanorder-of-magnitudehigher duetoitshigherprevalenceinuserqueries. We emphasize that observational datasets such as WildChat are subject to sampling bias and validity limitations. Our goal is not precise risk estimation, but to demonstrate a robust qualitative phenomenon: incorporating even coarse empirical prevalence can fundamentally alter the interpretation of safety benchmarks. This reframing preserves model rankings while extending evaluation beyond comparison toward understanding the relative types and magnitudes of risks models may expose in deployment. Relationship to Measurement Theory. This calibration exercise illuminates how benchmark scores relate to core measurement properties. Proportionality asks whether score differences correspond to proportional changes in risk. Within a single category, calibrated frequencies do scale proportionally with failure rate: doubling the failure rate doubles expected harmful exposures. 32 How Should AI Safety Benchmarks Benchmark Safety? However, this frequency-based proportionality does not extend to severity. Twice the failure rate on extremism prompts may produce far more than twice the real-world harm, as certain risks propagate nonlinearly through social systems. Invariance concerns whether metrics behave consistently across deployment contexts. Fig. 4 demonstrates that raw benchmark scores violate this property: identical refusal rates yield vastly different real-world exposure depending on category prevalence. A 0.9 refusal rate for sexual content and violence/extremism appears equivalent in benchmark terms, yet the calibrated frequencies differ by an order-of-magnitude. This suggests two complementary remedies: benchmarks could sample prompts proportionally to real-world prevalence, making aggregate scores more interpretable, or benchmarks could report calibrated frequencies alongside raw rates, making context-dependence explicit. Traceable calibration asks which real-world quantities scores approximate. Raw refusal rates approximate model behavior under a fixed prompt distribution, but this quantity is difficult to interpret in deployment terms. By multiplying failure rates by in-the-wild prevalence, we establish an explicit link between benchmark outputs and a concrete real-world quantity: expected frequency of harmful content exposure per query. This transformation exemplifies traceable calibration, converting abstract scores into deployment-grounded estimates. C.2. Translating Safety Benchmarks to Deployment Risk: A Fermi Approach The previous subsection demonstrated how calibrating benchmark frequencies against in-the-wild prevalence transforms model-centric scores into exposure estimates. However, exposure frequency alone does not capture risk: a high-frequency, low-severity harm may matter less than a rare but catastrophic one. Here, we extend this framework by tracing how prevalence propagates through subsequent stages of real-world impact, combining with severity to estimate the magnitude of losses rather than occurrence frequencies alone. In the tradition of Fermi estimation in physics and engineering (Fermi, 1945; Weinstein and Adam, 2008; Mahajan, 2010), we construct an illustrative calculation to demonstrate how benchmark failure rates might translate to deployment risk under plausible assumptions. Fermi estimation, which involves making approximate calculations with minimal data to achieve order-of-magnitude understanding, is standard practice for exploring system behavior when comprehensive empirical data remains unavailable (von Baeyer, 1993). As with scenario analysis more broadly, our goal is not predictive precision but structural clarity (Jones et al., 2001; Kosow and Gaßner, 2008): at what order-of-magnitude does the relationship between benchmark scores and real-world risk operate? Benchmark coverage versus real-world harm. Safety benchmarks stress contemporary models and rank systems via aggregate success or failure rates across behavior classes. While such coverage spans domains including misinformation, copyright, biological and chemical risks, and harassment, benchmarks typically do not provide a direct mapping from evaluation outcomes to realized real-world harm. We propose that translating benchmark performance to deployment risk may benefit from a multiplicative decomposition: Calibrated Risk =E[Harm] = B |z Benchmark Failure Rate × C |z Prevalence (composite) × S |z Severity whereBdenotes the benchmark failure rate (e.g., attack success rate),Crepresents a composite prevalence term capturing multiple real-world discount factors, andSquantifies the severity per realized harm. Current safety benchmarks report only B; researchers must estimateCandSfrom deployment data and domain-specific assessments. Consistent with scenario analysis traditions in probabilistic risk assessment (Kaplan and Garrick, 1981), this construction is explicitly illustrative; deployment contexts, user populations, and institutional safeguards vary substantially. Illustrative case: copyright-related behaviors.We demonstrate this framework using copyright violations as a worked example. To support transparency and reproducibility, the code and data file used for this analysis are available at https://anonymous.4open.science/r/ai-safety-benchmark. To make the pathway from benchmark failure to realized economic harm explicit, we decompose the composite prevalence termCinto a product of factors corresponding to distinct stages of deployment: C = C jailbreak × C enforce × C users × C queries , 33 How Should AI Safety Benchmarks Benchmark Safety? where each term reflects an empirically grounded filter from model behavior to realized liability. Benchmark failure rate (B ≈ 10 −2 ). HarmBench (Mazeika et al., 2024) reports that GPT-4 Turbo exhibits an attack success rate of approximately0.6%on copyright-related behaviors under the “Human Jailbreak” setting. This setting evaluates model behavior under a fixed set of in-the-wild jailbreak templates (Shen et al., 2024b), into which copyright-related behavior strings are inserted as user requests. We treat this quantity as a conditional failure probability under a specific adversarial prompt distribution. Jailbreak prevalence (C jailbreak ≈ 10 −2 ). WildChat (Zhao et al., 2024b) shows that prominent adversarial jailbreak prompt patterns appear in approximately1%of real-world user interactions (9,845 out of 1,039,785 queries). Using this aggregate jailbreak prevalence as a proxy for copyright-specific adversarial attempts likely overestimates copyright risk, since many jailbreaks target unrelated behaviors, while simultaneously underestimating it insofar as copyright-violating requests can be non-adversarial (e.g., direct requests to reproduce paywalled text). For our purposes, this proxy suffices to illustrate uncertainty propagation at the order-of-magnitude level. Enforcement rate (C enforce ≈ 10 −3 ). A critical ecological consideration is that infringement does not imply litigation or payment. TRAC reports 3,944 new federal copyright infringement cases filed annually (Transactional Records Access Clearinghouse (TRAC), 2016), while the Lumen Database receives approximately 5,000–7,000 takedown notices per day—the majority of which are DMCA (copyright-based) notices (Holland et al., 2019). Annualizing the latter yields roughly 2× 10 6 notices per year, giving: C enforce ≈ 3,944 2× 10 6 ≈ 2× 10 −3 This ratio captures the fact that the overwhelming majority of alleged infringements do not escalate to litigation or monetary liability. Platform scale (C users × C queries_per_year ≈ 10 6.5 ). We consider a medium-sized platform with10 3 daily active users and approximately10 1 queries per user per day, consistent with observed usage patterns for conversational AI systems (NerdyNav, 2025). This yields roughly 10 4 queries per day, or 3.65× 10 6 ≈ 10 6.5 queries per year. Severity (S ≈ 10 2.5 –10 6.5 ). The economic severity of copyright violations spans several orders of magnitude. Fig. 5 shows the distribution of U.S. copyright statutory damages awards from 2011–2020 (n=202) using the NYU Brady–Germano– Sprigman dataset (Brady et al., 2025). In log-space, representative cutpoints are: S μ−2σ ≈ $345≈ 10 2.5 , S μ ≈ $31,591≈ 10 4.5 , S μ+2σ ≈ $2.9× 10 6 ≈ 10 6.5 . Figure 5. Statutory Damages Distribution with Case Study Risk Assessment. Four risk levels defined by cutpoints atμ± 2σoflog 10 awards, U.S. copyright cases, 2011–2020 (n = 202). 34 How Should AI Safety Benchmarks Benchmark Safety? Combining these factors yields an illustrative expected annual liability: E[Annual Liability] = B× C jailbreak × C enforce × (C users × C queries_per_year )× S. Using the representative severity level S μ ≈ 10 4.5 , we obtain E[Annual Liability]≈ (10 −2 )× (10 −2 )× (10 −3 )× (10 6.5 )× (10 4.5 ) ≈ 10 4 USD (i.e., on the order of $10 4 per year). As shown in Fig. 5, this estimate is placed in empirical context. The blue star marks the illustrative annual liability implied by the benchmark calculation, positioned at the mean (μ) of the log-transformed statutory damages distribution. This placement reflects our intent to estimate baseline exposure under typical severity conditions, rather than a worst-case scenario driven by rare, catastrophic awards in the tail. This highlights that even a modest benchmark failure rateB, when propagated through realistic prevalence, enforcement, and scale factors, can generate nontrivial expected liability. This calculation demonstrates the structure of risk translation rather than deployment-ready estimates. The specific values depend heavily on application context (customer-facing chatbot versus internal tool), user population characteristics, organizational risk tolerance, and jurisdictional legal frameworks. We encourage practitioners to substitute domain-specific values while preserving the methodological framework that helps make implicit assumptions explicit. Discussion: Reconciling Risk Modeling with AI System Complexity.Risk modeling in AI safety remains a subject of active debate. In the context of system safety, Dobbe (2022) argues that Probabilistic Risk Assessment (PRA), while standard in some engineering domains, may be inappropriate for AI systems where failures emerge from complex sociotechnical interactions rather than component-level stochasticity. This suggests that even well-calibrated benchmark frequencies may provide false assurance if they obscure the constraint-satisfaction structure underlying real-world safety. Touzet et al. (2025) also emphasizes that safety cases cannot be a substitute for risk modeling, as they do not aim to estimate the likelihood and severity of risks. Despite these tensions, recent work continues to operationalize "benchmark risk" using likelihood–severity decompositions (i.e.,Risk = Probability× Severity) to conduct risk mitigation calculations in LLM evaluation (McGregor et al., 2025). Meanwhile, Wisakanto et al. (2025) formalizes risk level, mapping likelihood levels (ranging from10 −1 to10 −12 ) against harm severity (ranging from minor incidents to catastrophic outcomes, such as 1 to 500M+ fatalities). We emphasize that applying risk modeling framework is neither a complete nor a fully faithful model of safety. This limitation is well-recognized even in mature safety-critical domains such as nuclear power and aviation (Niehaus, 2002; Apostolakis, 2004). In those fields, risk modeling with PRA is explicitly treated as approximate, incomplete, and sensitive to "unknown unknowns" and sociotechnical dynamics (Zio, 2009). Therefore, critiques arguing that PRA is inappropriate for AI due to the system’s complexity do not uniquely apply to AI; they apply to all complex engineered systems embedded in sociotechnical environments (Hettinger et al., 2015) In practice, safety engineering resolves this tension not by abandoning risk modeling, but by using probabilistic estimates as decision-support tools: they serve to compare orders of magnitude, identify dominant risk contributors, and establish conservative lower bounds that are then augmented by safety factors and qualitative analysis (Cetiner et al., 2014; Zio, 2009). Our use of prevalence calibration and order-of-magnitude estimation follows this engineering tradition. We do not propose a universal mapping from benchmark scores to real-world harm, nor do we claim that probabilistic estimates capture the full dimensionality of AI safety. Rather, our goal is to make implicit assumptions explicit and demonstrate how benchmark results can be situated within a broader risk-reasoning workflow. This approach complements system-theoretic methods—such as Failure mode and effects analysis (FMEA) or System-theoretic process analysis (STPA)—which uncover hazards arising from interactions between human users and AI components that model-level evaluations often miss (Li and Chignell, 2022; Mylius, 2025). Ultimately, we aim to move beyond high-level frameworks by providing a quantitative, order-of-magnitude lens that accounts for component interactions and systematic effects without over-relying on exact point estimates. C.3. Community-Grounded Judgments Reduce Proxy–Impact Distance: Qualitative Example To illustrate why involving affected communities can reduce the distance between benchmark proxies and real-world harms, we draw on qualitative prompt–response examples from Ali et al. (2025) and the accompanying finding that 35 How Should AI Safety Benchmarks Benchmark Safety? rater disagreement is pervasive and systematically structured by demographic group. Despite sharing identical scenarios, demographic groups systematically rated the harms differently across dimensions (e.g., male participants rated responses less toxic than female participants; conservative and Black participants rated higher emotional awareness than liberal and White participants), implying that a single benchmark label can implicitly encode whose values define harm. Therefore, disagreement is often signal rather than annotator noise; collapsing to majority vote can erase minority viewpoints and mismeasure harms for those most exposed. These observations motivate incorporating feedback from affected communities and, where appropriate, preserving disagreement, because doing so better aligns scenario selection and risk operationalization with the harms people actually experience. D. Checklist Instrument and Worked Example The checklist below operationalizes our recommendations (R1–R10) into actionable items for benchmark design and reporting. We also provide multiple formats (spreadsheet, markdown) athttps://anonymous.4open.science/ r/ai-safety-benchmark/ for adoption and adaptation by researchers and practitioners. D.1. Checklist for Safety Benchmark Design Table 2. Checklist for safety benchmark design, aligned with Recommendations R1–R10. ModuleRecommendationChecklist Item Construct Coverage R1. Documenting Known Blind Spots a.State which risk types are evaluated (e.g., toxicity, bias, jailbreaks, misinformation) and how each is measured. b.Document excluded risks and deployment assumptions (e.g., single- turn only, English text only, no tool-use or multi-agent scenarios). c.Acknowledge that passing the benchmark does not guarantee safety against untested or emergent risks. d.Compare coverage against prior benchmarks and document differ- ences in risk categories or evaluation methods. R2.Expanding Known Boundaries a.Describe mechanisms for discovering novel risks (e.g., fuzzing, red-teaming, LM-generated evaluation). b. Include infrastructure for community contribution (e.g., submission portals with versioned integration and contributor credit). c.Describe update mechanisms: how and when new risks will be in- corporated (e.g., annual taxonomy review, continuous calibration). d.Incorporate multi-agent or interactive evaluation where emergent risks may arise through interaction. R3. Reframing ML Phe- nomena as Safety Con- cerns a.Address distribution shift as a safety-critical issue, not merely a performance limitation. b.Document potential annotation bias (annotator selection, disagree- ment patterns, demographic skew). c.Implement contamination detection and track temporal validity of evaluation sets. Risk Quantification R4.Calibrating Bench- mark Frequencies to Expo- sure Estimates a. Report results as “observed rate” or “failure frequency”; reserve “probability” for calibrated estimates. b. Calibrate benchmark rates against in-the-wild prevalence to support risk-relevant interpretation. c.Specify factors limiting generalization (e.g., limited prompt diver- sity, synthetic scenarios, missing user context). (Continued on next page) 36 How Should AI Safety Benchmarks Benchmark Safety? (Continued from previous page) ModuleRecommendationChecklist Item R5. Grounding Severity in Principled Frameworks a. Justify severity scales by citing sources (e.g., prior research, regula- tory standards, domain-specific frameworks). b.Clarify scale semantics: equal intervals, power-law relationships, or catastrophic thresholds. c. Reference established practices where applicable. R6. Accounting for Uncer- tainty Quantification a.Report confidence intervals, standard errors, or worst-case bounds for main metrics. b. Report inter-rater reliability (e.g., Cohen’sκ) for human or LLM- as-judge evaluation. c.Test robustness: report score variance across random seeds, prompt rephrasings, and evaluator versions. d. Apply explicit safety margins when extrapolating to deployment risk; note that real-world risk depends on user behavior, system safeguards, and context. Measurement Validity R7. Standardizing Safety Constructs with Trans- parency a.Specify harm constructs targeted (e.g., toxicity, bias, manipulation) and provide operational definitions for each. b. State whose values inform judgments of harm (expert assessments, policy frameworks, affected communities). c.Acknowledge contested normative choices in construct definitions. d.Articulate the relationship between measured proxies and real-world safety concerns. R8. Locking and Version- ing for Reproducibility a. Record model access details: interface (API or UI), access date, inference parameters (temperature, top-p), system prompt. b.Fix and report random seeds for data sampling, model inference, and evaluation. c. Version all evaluation components: LLM judge model, scoring rubric, annotation guidelines. d. Provide documentation (e.g., code repository, README) to enable independent replication. R9. Anchoring Proxies in Deployment Contexts a.State prompt sources: real user logs, researcher-designed, LLM- generated, or crowdsourced. Report validation if synthetic. b.Justify why test scenarios represent real-world use; if abstract or synthetic, state limitations. c.Include multi-turn evaluation if risks emerge over extended interac- tion (e.g., manipulation, trust exploitation). d. Document trade-offs when in-the-wild collection is constrained by privacy or curation opacity. R10. Iterative Refinement via Community Input a.Treat benchmarks as evolving instruments with risk-sensitive update cycles. b.Involve affected communities to assess whether scenarios reflect harms they experience. c. State whether the benchmark evaluates the model alone or within system context (user interaction, interface, safeguards). d. Acknowledge model-level scores̸=system-level safety; identify additional evaluation needed for deployment (e.g., user studies, sandbox simulations, post-deployment monitoring). 37 How Should AI Safety Benchmarks Benchmark Safety? D.2. Worked Example: AIR 2024 Checklist Assessment We next apply the checklist as a worked example by auditing AIR 2024 (Zeng et al., 2024). Assessment criteria.✓(addressed) indicates the benchmark explicitly and fully implements the checklist item with clear documentation;◦(partially addressed) indicates the item is mentioned or implemented incompletely, without full specification or justification;✗(not addressed) indicates no evidence the item was considered in the benchmark design or documentation. Table 3. Checklist assessment for AIR 2024, aligned with Recommendations R1–R10. Status indicators:✓= addressed,◦= partially addressed,✗ = not addressed. ModuleRecommendationChecklist ItemStatus Construct Coverage R1. Documenting Known Blind Spots a. State which risk types are evaluated and how each is measured.✓ b. Document excluded risks and deployment assumptions.◦ c. Acknowledge that passing does not guarantee safety against untested risks. ◦ d. Compare coverage against prior benchmarks.✓ R2. Expanding Known Boundaries a. Describe mechanisms for discovering novel risks.✗ b. Include infrastructure for community contribution.✗ c. Describe update mechanisms for incorporating new risks.◦ d. Incorporate multi-agent or interactive evaluation.✗ R3. Reframing ML Phenomena a. Address distribution shift as safety-critical.✗ b. Document potential annotation bias.✗ c. Implement contamination detection and track temporal validity.✗ Risk Quantification R4. Calibrating Benchmark Frequencies to Exposure Estimates a. Report results as “observed rate”; reserve “probability” for calibrated estimates. ◦ b. Calibrate benchmark rates against in-the-wild prevalence.✗ c. Specify factors limiting generalization.◦ R5.Grounding Severity in Principled Frameworks a. Justify severity scales by citing sources.◦ b. Clarify scale semantics: intervals, power-law, or thresholds.✗ c. Reference established practices.✗ R6. Systematic Uncertainty Quantification a. Report confidence intervals or standard errors.✗ b. Report inter-rater reliability.✓ c. Test robustness across seeds, prompts, evaluator versions.◦ d. Apply explicit safety margins when extrapolating to deployment.✗ Measurement Validity R7. Standardizing Safety Constructs with Transparency a. Specify harm constructs and provide operational definitions.✓ b. State whose values inform judgments of harm.✓ c. Acknowledge contested normative choices.◦ d. Articulate relationship between proxies and real-world safety.◦ R8. Locking and Versioning for Reproducibility a. Record model access details and inference parameters.◦ b. Fix and report random seeds.✗ c. Version all evaluation components.◦ d. Provide documentation for independent replication.✓ R9. Anchoring Proxies in Deployment a. State prompt sources and report validation if synthetic.✓ b. Justify why scenarios represent real-world use.◦ c. Include multi-turn evaluation for extended interaction risks.✗ d. Document trade-offs for in-the-wild collection constraints.✗ R10. Iterative Refinement via Community a. Treat benchmarks as evolving instruments.◦ b. Involve affected communities in assessment.✗ c. State whether benchmark evaluates model alone or within system context. ◦ d. Acknowledge model-level̸= system-level safety.◦ Summary Of the 37 checklist items, AIR 2024 fully addresses 7 (✓), partially addresses 15 (◦), and does not address 15 (✗). The benchmark demonstrates strength in construct coverage documentation (R1) and value specification (R7), but shows gaps in dynamic risk discovery (R2), ML phenomena as safety concerns (R3), uncertainty quantification (R6), and community engagement (R10). E. Coding Details Coding Dimensions We coded all surveyed benchmarks along multiple dimensions, including their assignment to Rumsfeld-style risk categories, whether benchmarks explicitly specify the risks not covered in their work, their reliance on fixed versus dynamic data, the use of binary outcome metrics, distinctions between harm severity levels and whether 38 How Should AI Safety Benchmarks Benchmark Safety? such distinctions are theoretically grounded, explicit treatment of uncertainty, grounding of proxies in external standards or regulations, and whether evaluations are limited to single-turn interactions. Coding protocol provided in additional file. Table 4. Coding results for all surveyed benchmarks. One author performed initial coding; a second author reviewed all assignments, with disagreements resolved through discussion. Column definitions: Rumsfeld: Rumsfeld category assignment (K = Known knowns, KU = Known unknowns, UK = Unknown knowns, U = Unknown unknowns). Uncovered Doc.: Whether the benchmark explicitly specifies the risks it uncovers, including unexpected, unknown, or unforeseen risks. Fixed Data: Whether the benchmark relies on a fixed, static dataset. Binary Metric: Whether the benchmark reduces evaluation to binary outcome proportions (e.g., harmful/safe, reject/not reject, biased/unbiased, attack success/failure) as the primary metric. Sev. Level: Whether the benchmark distinguishes between different levels of harm severity (e.g., low/medium/high or Level 1–5). Sev. Grd.: If severity levels are distinguished, whether they are grounded in an explicit theoretical framework or external standard (e.g., regulatory guidance or established harm taxonomies). Uncert.: Whether the benchmark explicitly identifies sources of uncertainty or variation (e.g., model sensitivity, response stochasticity, sampling variation, or evaluator disagreement). Proxy Grd.: Whether benchmark definitions or evaluation proxies are explicitly grounded in established frameworks, regulations, policies, or societal standards. Single-turn: Whether the benchmark evaluates models in isolation via a single question-answer interaction, without contextual information or multiple trials. Notation.✓denotes yes,◦denotes partial or mixed support,✗ denotes no, and NA indicates not applicable. RumsfeldUncovered Doc.Fixed DataBinary MetricSev. LevelSev. Grd.Uncert.Proxy Grd.Single-turn Benchmark CATQA (Bhardwaj et al., 2024)K◦✓✗NA✓ HOLISTICBIAS (Smith et al., 2022)K✓◦✗NA✓ SAFETEXT (Levy et al., 2022)K◦✓◦✗NA✓✗✓ WALLEDEVAL (Gupta et al., 2024b)K◦✓✗NA✓ ◦✓ TOXIGEN (Hartvigsen et al., 2022)K✓◦✓✗✓ JADE (Zhang et al., 2023)KU◦✗◦✗NA✓ ◦✓ PROSOCIALDIALOG (Kim et al., 2022)K◦✓✗✓✗ StrongREJECT (Souly et al., 2024)K✓✗✓✗✓ XSTEST (Röttger et al., 2023)K◦✓✗NA✓✗✓ Cognitive Biases (Malberg et al., 2024)K◦✓✗NA✓ Non-Discrimination (Zhang et al., 2018)K✗✓✗NA✓✗✓ Societal Bias VLMs (Sathe et al., 2024)K◦✓✗NA✓ AART (Radharapu et al., 2023)KU✓✗✓✗NA✓ AAVENUE (Gupta et al., 2024c)K◦✓◦✗NA✓✗✓ Ego-View Accident (Fang et al., 2024)K◦✓✗NA✓ ◦✓ Adversarial VQA (Li et al., 2021)KU✗✓✗NA✓✗✓ Adversarial GLUE (Wang et al., 2021)UK✗✓✗NA✓✗✓ SAFETY-TUNED LLAMAS (Bianchi et al., 2023) K✗✓✗✓ ◦✓ AgentDojo (Debenedetti et al., 2024)KU✓✗✓✗NA ◦✗ AILUMINATE (Ghosh et al., 2025)K✓✗NA✓ AIR-BENCH 2024 (Zeng et al., 2024)K◦✓◦✓ ALERT (Tedeschi et al., 2024)K◦✓✗NA✓✗✓ QA-LIGN (Dineen et al., 2025)K✓◦✓✗✓✗✓ Wang et al. (2024a)K◦✓✗NA✓✗✓ Sotnikova et al. (2021)K✓◦✓✗✓ Alfadel et al. (2020)K✓NA C2SaferRust (Nitin et al., 2025)K✓◦✗NA✓✗ Huang et al. (2022)K✓✗NA✓✗ (Wan et al., 2023)K✗✓◦✗NA✓✗✓ ArtPrompt (Jiang et al., 2024b)KU✓✗◦✓ ◦✓ WildTeaming (Jiang et al., 2024a)KU✓✗NA✓✗✓ Huang et al. (2023)KU◦✓✗NA✓ ◦✓ Athena (Sadhu et al., 2024)K✓✗✓✗ BackdoorLLM (Li et al., 2024b)KU◦✓✗NA✓ ◦✗ BBG (Jin et al., 2025a)K◦✓✗NA✓✗ BBQ (Parrish et al., 2021)K✓✗NA✓ BeaverTails (Ji et al., 2023)K✗✓◦✗NA✓✗✓ KOFFVQA (Kim and Jung, 2025)K✗✓✗✓✗✓✗✓ Wang et al. (2025a)KU✗✓✗NA✓✗✓ Flames (Huang et al., 2024a)K✗✓◦✓ Liang et al. (2023)K✗✓✗NA✓✗✓ Verma and Bharadwaj (2025)K✗✓✗NA✓✗✓ Laszkiewicz et al. (2024)K✓◦✗NA✓ SAFE (Yu et al., 2024c)K✗✓◦✗NA✓✗✓ Kirk et al. (2021)K✗✓✗NA✓✗✓ BOLD (Dhamala et al., 2021)K✓◦✗NA✓✗✓ Xu et al. (2021)K✗✓✗✓✗ Jaiswal et al. (2024)K✗✓✗NA✓ Dinan et al. (2019)KU✗✓✗NA✓✗ Bjørgen et al. (2018)K◦✗✓✗NA✓✗NA Continued on next page 39 How Should AI Safety Benchmarks Benchmark Safety? Table 4 – Continued from previous page RumsfeldUncovered Doc.Fixed DataBinary MetricSev. LevelSev. Grd.Uncert.Proxy Grd.Single-turn Benchmark CALM (Gupta et al., 2024a)K✓✗NA✓ Chen et al. (2023)K◦✓◦✗NA✓ ◦✗ RuLES (Mu et al., 2024)K✓✗NA✓ ◦✗ CARNOVEL (Filos et al., 2020)UK◦✓◦✗NA✓✗ Islam et al. (2021)K✗✓✗NA✓NA OccuGender (Chen et al., 2025)K✗✓✗NA✓✗ CDEval (Wang et al., 2024b)K✗✓◦✗NA✓ CHBias (Zhao et al., 2023)K✗✓✗NA✓ ◦✗ CHiSafetyBench (Zhang et al., 2024a)K✗✓◦✓✗ CIF-Bench (Li et al., 2024c)K✗✓◦✗NA✓✗✓ CBBQ (Huang and Xiong, 2024)K✗✓NA✓ OW-DFA (Zheng et al., 2025)KU✓✗✓✗NA✓✗✓ CPO (Guo et al., 2024)K◦✓✗✓✗✓✗ CoSafe (Yu et al., 2024b)K✗✓✗NA✓✗ CrowS-Pairs (Nangia et al., 2020)K◦✓✗NA✓ CVE-Bench (Wang et al., 2025b)K✗✓✗NA✓✗ Cybench (Zhang et al., 2025b)K◦✓◦✗NA ◦✓✗ Sambasivan et al. (2021)K◦✗NA✓✗ DELPHI (Sun et al., 2023a)K✗✓◦✗NA✓✗✓ DICES (Aroyo et al., 2023)K◦✓◦✓✗✓✗ Perez et al. (2023)U✓◦✗NA✓ discrim-eval (Tamkin et al., 2023)K◦✓◦✗NA✓ Hall et al. (2023)K◦✓✗NA✓✗✓ DiversityMedQA (Rawat et al., 2024)K✗✓✗NA✓ JailbreakHub (Shen et al., 2024a)KU◦✓✗ Do-Not-Answer (Wang et al., 2024c)K✓◦✓✗✓✗✓ MACHIAVELLI (Pan et al., 2023)KU✓✗✓✗ Liu et al. (2020)K✗✓◦✗NA✓✗✓ ConfAIde (Mireshghallah et al., 2024)K✓◦✗NA✗✓✗ LabellessFace (Ohki et al., 2024)K✓✗NA✓✗✓ ePiC (Ghosh and Srivastava, 2022)K✓✗NA✓✗✓ GenMO (Bajaj et al., 2024)K✗✓✗NA✓✗✓ RealToxicityPrompts (Gehman et al., 2020) K✓◦✗NA✓✗✓ Tri-HE (Wu et al., 2025)K✗✓✗NA✓✗✓ Li et al. (2024d)KU◦✓◦✗NA✓✗✓ Kiritchenko and Mohammad (2018)K✓✗NA✓ ◦✓ AdvMark (Wu et al., 2024)KU✗✓◦✗NA✓ FACET (Gustafson et al., 2023)K◦✓✗NA✓ FairLex (Chalkidis et al., 2022)K◦✓✗NA✓ Siddiqui et al. (2022)K✗✓◦✗NA✗✓ FFT (Cui et al., 2023)K✗✓◦✗NA✓ ◦✓ Filipino (Gamboa and Lee, 2025)K◦✓✗NA✓ ◦✓ FLEX (Jung et al., 2025)KU◦✓✗NA✓✗✓ FLAIR (Song et al., 2022)K✓✗NA✓✗✓ FrenchCrowS-Pairs (Névéol et al., 2022)K◦✓✗NA✓ ◦✓ FrenchToxicityPrompts(Brunand Nikoulina, 2024) K✓◦✓✗✓✗✓ Z’eroe (Eger and Benz, 2020)K◦✓◦✗NA✓✗✓ WinoBias (Zhao et al., 2018)K✗✓✗NA✓ GeoNet (Kalluri et al., 2023)K✓✗NA✓✗✓ CPAD (Liu et al., 2023)K✗✓✗NA✓ ◦✗ GPTFuzz (Yu et al., 2024a)KU◦✗✓✗NA✓✗✓ CoP (Xiong et al., 2025)KU✓✗✓✗✓ ◦✗ SafeWatch-Bench (Chen et al., 2024b)K✓◦✗NA✓✗ HarmBench (Mazeika et al., 2024)KU✓✗NA✓✗ HRS-Bench (Bakr et al., 2023)K✗✓✗NA✓ ◦✓ FVQA2.0 (Lin et al., 2023)KU◦✓✗NA✓ ◦✓ HypoTermQA (Uluoglakci and Temizel, 2024) KU◦✓✗NA✓✗✓ IHEval (Zhang et al., 2025c)K✓◦✗NA ◦✓✗ IndiBias (Sahoo et al., 2024)K◦✓◦✗NA✓ ◦✓ UNQOVER (Li et al., 2020a)K◦✓✗NA✓ ◦✓ InjecAgent (Zhan et al., 2024)K◦✓✗NA✓✗ ADVQA (Sung et al., 2025)KU◦✓✗NA✓ JailbreakBench (Chao et al., 2024)K◦✗✓✗NA✓✗ Al Ghanim et al. (2024)K✓✗NA✓✗ JailBreakV-28K (Luo et al., 2024)K✗✓✗NA✗✓ JobFair (Wang et al., 2024d)K✓◦✓✗✓✗ KoBBQ (Jin et al., 2024)K◦✓✗NA✓ KorNAT (Lee et al., 2024)K◦✓◦✗NA✓ LatentJailbreak (Qiu et al., 2023)K◦✓✗NA✓✗ Continued on next page 40 How Should AI Safety Benchmarks Benchmark Safety? Table 4 – Continued from previous page RumsfeldUncovered Doc.Fixed DataBinary MetricSev. LevelSev. Grd.Uncert.Proxy Grd.Single-turn Benchmark CCLR (Goodier and Campbell, 2023)UK✗✓✗NA✗NA Jain and Krishnamurthy (2025)K◦✗✓ ◦✗ LLMArena (Chen et al., 2024a)U◦✗NA✓✗ SAD (Laine et al., 2024)K✓✗NA✓ ◦✓ MedHALT (Pal et al., 2023)K◦✓✗NA✓ MEDFAIR (Zong et al., 2023)K◦✓✗NA✓✗✓ MedSafetyBench (Han et al., 2024)K◦✓✗✓ Ahn and Oh (2021)K◦✓✗NA✓ ◦✓ MLeVLM (Xu et al., 2024)K◦✓◦✓NA✓ ◦✗ MMEvalPro (Huang et al., 2025)K◦✓✗NA✓✗ ModSCAN (Jiang et al., 2024c)K✓◦✗NA✓ MultiRobustBench (Dai et al., 2023)K✓◦✗NA✓✗✓ Rudinger et al. (2018)K◦✓✗NA✓ OKTest (Shi et al., 2024)K✓✗NA✓✗✓ CLCA (Liu et al., 2025a)K✓◦✗NA✓✗ Shaikh et al. (2023)K✗✓✗NA✓✗ Sun et al. (2022)K✓✗NA✓✗ Vashishtha et al. (2023)K◦✓✗NA✓✗✓ CAMeL-2 (Naous and Xu, 2025)K✓✗NA✓ ◦✓ Chatterjee et al. (2024)K◦✓✗NA✓✗✓ OR-Bench (Cui et al., 2025)KU◦✗✓✗NA✓ PAPILLON (Li et al., 2025)K✓✗NA✓✗✓ PERSONA (Castricato et al., 2025)K✓✗NA✓ PLUE (Chi et al., 2023)K◦✓✗NA✓ ◦✓ Poser (Clymer et al., 2024)K◦✓✗NA✓ PrivLM-Bench (Li et al., 2024e)K✗✓◦✗NA✓ PROMPTEVALS (Vir et al., 2025)K◦✓✗NA✓ MLLMU-Bench (Liu et al., 2025b)K✓◦✓NA✗✓ RADDLE (Peng et al., 2021)KU✗✓✗NA✓ ◦✗ RAGTruth (Niu et al., 2024)K◦✓NA✓✗ Bhatt et al. (2022)K✓✗NA✓ ◦✓ MDIT-Bench (Jin et al., 2025b)K✓✗NA✓✗ REAP (Hingun et al., 2023)K✗✓◦✗NA✓ Ganguli et al. (2022)K✓NA◦✓NA✓ ◦✗ RED-EVAL (Bhardwaj and Poria, 2023)KU◦✓✗NA✓ ◦✗ RedditBias (Barikeri et al., 2021)K◦✓✗NA✓✗ DeMET (Levy et al., 2024)K✗✓◦✗NA✓ TLDR (Fu et al., 2025)K✗✓◦✗NA✓ ◦✓ ROBBIE (Esiobu et al., 2023)K✓✗NA✓ ◦✓ Robo3D (Kong et al., 2023)K◦✓✗✓ ◦✓ GQA-OOD (Kervadec et al., 2021)K◦✓✗NA✓✗✓ S-Eval (Yuan et al., 2025)K◦✗✓✗NA✓✗ LaMaSafe (Wang et al., 2024e)K◦✓✗NA✓✗ SafeBench (Xu et al., 2022)KU✗◦✗NA ◦✓✗ Sun et al. (2023b)K◦✓✗NA✓✗✓ SafetyBench (Zhang et al., 2024b)K✓✗NA✓✗✓ SALAD-Bench (Li et al., 2024a)K◦✓✗NA✓✗ Sun and Miceli-Barone (2024)K✓◦✗NA✓✗✓ SG-Bench (Mou et al., 2024)K◦✓✗NA✓ ◦✗ SimpleSafetyTests (Vidgen et al., 2023)K✓✗NA✓ SoFa (Manerba et al., 2024)K✓✗NA✓ ◦✓ Sadhu et al. (2025)K◦✓✗NA✓ ◦✓ SORRY-Bench (Xie et al., 2025)K✓✗NA✓ StereoSet (Nadeem et al., 2021)K✓✗NA✓✗ BenchPress (Mitra et al., 2019)K✓✗NA✓✗ TensorTrust (Toyer et al., 2024)KU✓✗NA✓✗✓ COCONOT (Brahman et al., 2024)K◦✓✗NA✓ Aakanksha et al. (2024)K✓✗NA✓✗✓ SwedishWinogender (Hansson et al., 2021) K◦✓✗NA✓ WMDP (Li et al., 2024f)K◦✓✗NA✓ Sheng et al. (2019)K✗✓✗NA✓ ◦✓ Berman and Albright (2017)UK✓NA✗NA✗NA TheGreatestGood (Marraffini et al., 2024)K◦✓✗NA✓ TAB (Pilán et al., 2022)KU◦✓◦✓NA EmpatheticDialogues (Rashkin et al., 2019) K✗✓✗✓✗✓✗ Zhou et al. (2022)K✓✗NA✓✗ MMHB (Tan et al., 2024)K◦✓◦✗NA✗ ◦✓ Liang et al. (2021)K◦✓✗NA✓✗ Li et al. (2023)K✗✓✗NA✓✗✓ Santosh et al. (2024)K◦✓◦✗NA✓✗✓ Continued on next page 41 How Should AI Safety Benchmarks Benchmark Safety? Table 4 – Continued from previous page RumsfeldUncovered Doc.Fixed DataBinary MetricSev. LevelSev. Grd.Uncert.Proxy Grd.Single-turn Benchmark Bai et al. (2022)K✓✗✓✗NA✓✗ TrustLLM (Huang et al., 2024b)K✗✓◦✓✗ TruthfulQA (Lin et al., 2022)K◦✓◦✓✗✓ ◦✓ Qi et al. (2024)KU✓◦✓✗✓ CII-Bench (Zhang et al., 2025d)K◦✓◦✓✗✓✗✓ UNQOVER (Li et al., 2020b)K✓◦✗NA✓✗✓ Zhao et al. (2021)K✓✗NA✓ MHaluBench (Chen et al., 2024c)K✓✗NA✓✗✓ AdvBench (Zou et al., 2023)KU◦✓◦✗NA✓✗✓ Kour et al. (2023)KU◦✓✗NA✓ ◦✓ LMMs-Eval (Zhang et al., 2025a)UK◦✗NA✓✗ VALUE (Ziems et al., 2022)K✓◦✗NA✓✗✓ MixCuBe (Kim et al., 2025)K✗✓✗NA✓✗✓ B-score (Vo et al., 2025)K✗✓◦✗NA✓✗ WinoQueer (Felkner et al., 2023)K✓✗NA✓ SeqAR (Yang et al., 2025)KU◦✓✗NA✓✗✓ Wang (2024)KU◦✗✓✗NA✓✗✓ GEST (Pikuliak et al., 2024)K✓◦✗NA✓ TiEBe (Almeida et al., 2025)K✓✗NA✓✗✓ WorldCuisines (Winata et al., 2025)K◦✓◦✗NA✓✗✓ VITAL (Shetty et al., 2025)K✓◦✗NA✓ ◦✓ 42