Paper deep dive
Epistemic Norms for AI Safety and Alignment Research
Keivan Navaie
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment research has a different mission: to ensure that catastrophic failures never occur, under sparse evidence, adversarial dynamics, and fat-tailed risk. We argue that the two domains differ along two analytically independent axes---{\it capability profile}, demonstrating the absence of hazardous behaviours rather than the presence of positive capabilities, and {\it risk profile}, bounding worst-case outcomes under fat-tailed uncertainty rather than optimising average-case performance---and that mainstream epistemic practices are inadequate on both. Building on a structured synthesis grounded in a preregistered bibliometric baseline, we identify five cross-cutting gap dimensions in current alignment research, including the near-absence of institutionalised independent verification. To address these gaps, we propose {\sc ECAISA}, an Epistemic Code for AI Safety and Alignment comprising eight principles, a three-level scoring rubric, a four-level disclosure ladder that reconciles transparency with information-hazard and commercial-confidentiality constraints, a tiered applicability scheme, an information-hazard adjudication procedure, and seven anti-gaming mechanisms. {\sc ECAISA} does not certify that any AI system is safe; it constrains how safety-relevant research claims are documented, checked, and relied upon, with auditability rather than certification as its governance target.
Tags
Links
- Source: https://arxiv.org/abs/2607.24243v1
- Canonical: https://arxiv.org/abs/2607.24243v1
Trouble viewing inline? Open PDF directly â
Full Text
129,932 characters extracted from source content.
Expand or collapse full text
1 Epistemic Norms for AI Safety and Alignment Research Keivan Navaie Lancaster University, UK (k.navaie@lancaster.ac.uk) Abstract Mainstream AI research emphasises capability growth and tolerates low failure rates when average- case performance is high. AI safety and alignment research has a different mission: to ensure that catastrophic failures never occur, under sparse evidence, adversarial dynamics, and fat-tailed risk. We argue that the two domains differ along two analyHcally independent axes â capability profile (demonstraHng the absence of hazardous behaviours versus the presence of posiHve capabiliHes) and risk profile (bounding worst-case outcomes under fat-tailed uncertainty versus opHmising aver- age-case performance) â and that mainstream epistemic pracHces are inadequate on both. Building on a structured synthesis grounded in a preregistered bibliometric baseline, we idenHfy five cross- cuOng gap dimensions in current alignment research, including the near-absence of insHtuHonalised independent verificaHon. To address these gaps we propose ECAISA, an Epistemic Code for AI Safety and Alignment comprising eight principles, a three-level scoring rubric, a four-level disclosure ladder that reconciles transparency with informaHon-hazard and commercial-confidenHality constraints, a Hered applicability scheme, an infohazard adjudicaHon procedure, and seven anH-gaming mecha- nisms. ECAISA does not cerHfy that any AI system is safe; it constrains how safety-relevant research claims are documented, checked, and relied upon, with auditability rather than cer4fica4on as its governance target. A retrospecHve rubric audit (Îș = 0.79) demonstrates instrument feasibility; a four-stage validaHon roadmap is proposed. 1. Introduc0on Epistemic norms tell a research community when a claim is âknowledge-grade,â that is, jusHfied enough to count as reliable knowledge. Under classical reliabilist theories, a belief qualifies as knowledge only if the inferenHal process that produced it is truth-conducive and reliably leading from true premises to true conclusions (Armstrong 1973). Robert K. Merton arHculated this ideal in four operaHonal norms: communalism (open sharing of findings), universalism (method over authority), disinterestedness, and organised scepHcism (Merton 1942/1973). In pracHce, an epistemic norm is a concrete standard that tells researchers when to accept, share, or challenge a claim; examples include âfindings must be independently replicableâ, âclaims must survive adversarial challengeâ, and âhypotheses must be preregisteredâ. Such norms are the operaHonal counterpart to the abstract ideals of communalism and scepHcism. Mainstream AI has, to a meaningful degree, operaHonalised these epistemic ideals by anchoring progress to shared public benchmarks and leaderboards, insHtuHng reproducibility checklists and arHfact-review tracks, and normalizing code/data release alongside rigorous (if imperfect) peer review. ImageNet, GLUE, and SuperGLUE exemplify community benchmarks that make results publicly comparable (Russakovsky et al. 2015; Wang et al. 2019, 2020). NeurIPSâs reproducibility program and related efforts formalize reporHng standards and require code, data, and hyperparameters for audit (Pineau et al. 2021; Gundersen & Kjensmo 2018). And while incenHves 2 remain uneven, norms of releasing implementaHon details and subjecHng work to conference/journal peer review have helped sustain a baseline of credibility (Raff 2019; Langley 2019). AI safety and alignment research poses a fundamentally different goal and risk profile: its mission is prevenHve to ensure that certain harmful outcomes (e.g., loss of human control over advanced AI) never occur whereas mainstream AI primarily seeks to create capabiliHes and measures success by posiHve task performance. The concept of alignment as a prevenHve, risk centred- agenda separate from purely capability -driven AI development was first arHculated by Bostrom (2011, 2014) and later elaborated by Russell, Dewey & Tegmark (2015) and Amodei et al. (2016). We use âAI safetyâ to denote the broad sociotechnical effort to prevent and miHgate harms from AI systems across their lifecycle, spanning technical controls, organisaHonal processes, and governance standards such as the NIST AI RMF, ISO/IEC 42001, and the EU AI Act. By contrast, âAI alignmentâ names the technical problem of ensuring that highly capable systemsâ objecHves and behaviours remain acceptably compaHble with human values under worst-case, adversarial, and distribuHon- shin condiHons. Alignment therefore carries a prevenHve mission (âprove the negaHveâ) and a fat- tailed risk profile, requiring very low tolerance for catastrophic error and privileging adversarial verificaHon over average-case performance gains. Throughout, we treat alignment as a core but not exhausHve component of the wider AI safety agenda; when we say âAI safety and alignmentâ, we refer to this coupled scope. In scienHfic terms, alignment is an evidence-of-absence problem: we must jusHfy, to very high confidence, that no unacceptable failure will occur, not merely show a posiHve capability. Because many plausible harms (e.g., an accidental existenHal catastrophe) would be irreversible, the burden of proof is sharply asymmetric and precauHonary. The challenge is compounded by the possibility that advanced AI systems behave adversarially or decepHvely, strategically masking misalignment during evaluaHon, which makes some failures intrinsically hard to foresee. Evidence for such adversarial/decepHve dynamics includes work on decep4ve alignment (Hubinger et al. 2019), Alignment Faking (Greenblap et al. 2024) and adversarial examples (Goodfellow, Shlens & Szegedy 2015), with further catalogues of hard-to-anHcipate safety failures and goal mis-generalizaHon in Amodei et al. (2016) and Langosco et al. (2022). Because we cannot ethically âtest to failureâ when the downside could be civilisaHon-scale harm, alignment research must proceed under sparse, indirect evidence, deep epistemic uncertainty, and fat-tailed risk where worst-case losses dwarf average performance. Any claim of safety or alignment therefore demands evidence far stronger than the metrics that suffice for ordinary ML benchmarks evidence organised as a safety case rather than a performance plot. This stricter standard echoes a growing body of work: Shevlane et al. (2023) call for âvery strong assuranceâ in alignment evaluaHons; Leveson (2012) disHnguishes rigorous safety assurance from mere reliability; Amodei et al. (2016) advocate accident-oriented validaHon for AI systems; and Bloomfield and Rushby (2024) along with Hilton et al. (2025) adapt high-assurance safety-case methodology specifically to AI. Furthermore, in most AI fields a 99.9% success rate might be considered excellent; but in alignment, even a 0.1% risk of failure could hide an existenHal catastrophe. Standard AI research incenHves (novelty, rapid publicaHon) tend to undervalue the extreme tail risks of AI, since they reward average-case improvements. We therefore argue that AI safety research instead requires an epistemic regime of extreme rigour: full transparency of methods, systemaHc adversarial tesHng, and very conservaHve inference. This mirrors the ethos of high-reliability engineering domains such as aviaHon, medicine, nuclear power, and high-assurance cyber security, where deployment is authorised only aner stringent demonstraHons of safety or resilience. Commercial aviaHon, for example, treats catastrophic-failure 3 probabiliHes below 10â»âč per flight-hour as a design-assurance target (FAA Advisory Circular 25.1309- 1B); criHcal-infrastructure cyber security demands formal verificaHon, exhausHve penetraHon tesHng, and zero-day resilience before operaHonal release. The 10â»âč figure serves here as a moHvaHng analogy, not as a computaHonal target for alignment: alignment lacks the actuarial base-rate data, well-defined failure modes, and staHonary dynamics that make a specific numerical threshold well- posed. The principle we take from these precedents is evidence-proporHonality, namely that the stringency of the evidence required for a safety claim should scale with the severity and irreversibility of the claimâs potenHal failures, not a specific probability target. By this principle, every safety claim in AI alignment should be scruHnised as though it were false, and even seemingly negligible risks cannot be dismissed without jusHficaHon. In this paper, we adopt a mixed methodology that combines conceptual analysis with comparaHve benchmarking. We begin by idenHfying the key asymmetries between alignment and convenHonal AI research and then conduct a gap analysis to pinpoint where mainstream epistemic norms fall short for alignmentâs needs. We then draw on lessons from safety-criHcal fields and from the open science and metascience literature to inform this analysis. Based on these insights, we formulate a concrete proposal: a codified Epistemic Code for AI Safety and Alignment (ECAISA) consisHng of eight core principles tailored to the alignment community. Our core contribuHons are: (i) a gap analysis between mainstream AI research norms and alignment requirements, grounded in a preregistered bibliometric baseline (§3.1); (i) the ECAISA code of eight epistemic principles, supported by a three-level scoring rubric for applicaHon by venues, funders, and laboratories (§5, Appendix A); (i) a governance crosswalk that maps each principle to NIST AI RMF 1.0, ISO/IEC 42001, and the EU GPAI Code of PracHce, with explicit statement of what ECAISA adds (§6.1, Table 7); (iv) a retrospecHve rubric audit of eight alignment papers (Îș = 0.79) demonstraHng instrument feasibility, together with a four-stage validaHon roadmap (§7); and (v) an analysis of adopHon dynamics that treats unilateral adopHon as a coordinaHon problem and idenHfies three insHtuHonal paths past first-mover disadvantage (§7.2). In the remainder of this paper, we elaborate each of these contribuHons, beginning with the fundamental epistemic differences that set alignment apart. 2. Epistemic Asymmetries Between Mainstream AI and Alignment Two asymmetries disHnguish safety and alignment research from mainstream AI. We state each along a single axis, so that the two are analyHcally disHnct and the gap analysis of §3 does not merely restate them. The first is an asymmetry of capability profile: mainstream AI demonstrates what a system can do, whereas alignment must demonstrate what a system cannot be made to do. The second is an asymmetry of risk profile: mainstream AI opHmises average-case performance, whereas alignment must bound worst-case outcomes under fat-tailed, partly adversarial uncertainty. These two axes are related but independent; a study can be strong on one and weak on the other. 2.1 Capability profile: demonstra6ng absence versus presence Mainstream AI demonstrates posiHve capabiliHes: higher accuracy, more fluent generaHon, stronger reward. The object of demonstraHon is what the system can do. Alignment inverts this: the object of demonstraHon is what the system must not do, even when pushed. That is, alignment must provide evidence for the absence of hazardous behaviours, including behaviours that adversarial users, distribuHon shins, or the system itself may elicit under condiHons not present in the evaluaHon. Safety-criHcal engineering offers the relevant analogy. Avionics and medical-device cerHficaHon do not rely on strong mean performance; they require demonstraHons that specific failure modes will 4 not occur at unacceptable rates across operaHng condiHons (Liplewood and Strigini 1993; Leveson 2012; RTCA DO-178C 2011; ISO 14971:2019; FAA AC 25.1309-1B). The stance that follows is Popperian (Popper 1959; Mayo 2018): safety claims are treated as conjectures that must survive severe tests, not as summaries of observed successes. In pracHce this means systemaHc adversarial tesHng, scenario analysis, and formal verificaHon applied to candidate failure modes (Amodei et al. 2016; Hubinger et al. 2019; Shevlane et al. 2023; Madry et al. 2018; Katz et al. 2017). A 99.9% benchmark score is less informaHve than the 0.1% of cases that failed, because it is the failure set that the safety claim must rule out. This asymmetry is about the object of demonstraHon; it is orthogonal to how the probability of failure is distributed, which is the subject of the second asymmetry below. 2.2 Risk profile: worst-case under fat-tailed uncertainty Mainstream AI evaluaHon opHmises and reports average-case performance. Alignment must instead bound the probability and magnitude of worst-case outcomes under distribuHons that are fat-tailed, sparsely sampled, and partly adversarial. The disHncHon mapers because a system can post excellent average performance while concealing rare but severe failures (Cirillo and Taleb 2020; Bostrom 2011; Shevlane et al. 2023). Where the loss distribuHon has a long tail, mean performance is not a summary of the object of interest; the tail is. The analogy to high-reliability engineering is instrucHve here. Nuclear regulators do not judge safety by expected outcomes; they require defence in depth, redundancy, and conHnuous audiHng of the assumpHons behind the safety case, and they quanHfy residual risk using explicit uncertainty analysis rather than point esHmates (IAEA SSR-2/1 Rev.1, 2016; NRC Regulatory Guide 1.174 Rev.3, 2018; Liplewood and Strigini 1993; Leveson 2012; NRC NUREG-1855 Rev.1, 2017). Alignment inherits this posture in full, and adds one feature that high-reliability engineering largely does not face: the adversarial component of the risk profile. Advanced models may mask unsafe behaviour during evaluaHon, so an apparent pass may itself be evidence of decepHve alignment (Hubinger et al. 2019). This requires conHnuous red-teaming, scepHcism toward tests that appear to pass, and formal safety-assurance methodology (Bloomfield and Bishop 2010; ISO/IEC 15026-2:2011; Shevlane et al. 2023; Perez et al. 2022; Ganguli et al. 2022). The risk profile is about the probabilisHc target of evidence; it is independent of, and orthogonal to, the capability profile set out above. A study may opHmise for worst-case performance and sHll commit the mainstream-capability framing by reporHng only successes (weak on capability profile). Likewise, a study may adopt the absence-of-hazard framing but report only mean-adversarial performance (weak on risk profile). TreaHng the two as independent lets the gap analysis in §3 diagnose failure modes along one axis without covertly invoking the other. 2.3 Safety as a sociotechnical property of the research record A clarificaHon of scope is necessary before proceeding to the gap analysis. Safety is not a property of technical systems in isolaHon but of how those systems are developed, evaluated, and deployed within social and insHtuHonal contexts (Leveson 2012; Wieringa 2020). ECAISA operates at a level consistent with this framing. It specifies the evidenHary pracHces that make a research-phase safety claim auditable, that is, inspectable and contestable by independent reviewers, funders, oversight bodies, and affected communiHes. The object of ECAISAâs requirements is the research record: the set of artefacts, disclosures, and documented pracHces that consHtute the evidence base for a safety-relevant claim. This is itself a sociotechnical object: produced within insHtuHonal contexts, shaped by incenHve structures, and consumed by decision-makers who rely on it to authorise deployment, allocate funding, or set policy. The five epistemic gaps idenHfied in §3 are therefore gaps in the social producHon of evidence, not merely technical measurement failures; ECAISAâs principles are, in this sense, governance intervenHons as much as epistemic ones. 5 We also acknowledge the epistemological limits of any assurance framework for open-ended systems (Narayanan and Kapoor 2023). Because the input and output spaces of capable AI systems are effecHvely unbounded, no evaluaHon programme can close the denominator of the failure probability; comprehensive safety assurance for general-purpose systems is not achievable by any finite set of tests. ECAISA does not apempt to achieve it. The appropriate analogy is clinical research: randomised trials cannot enumerate all paHent contexts or eliminate residual risk, yet structured evidenHary standards demonstrably improve the contestability and reliability of therapeuHc claims relaHve to unstructured pracHce. ECAISA makes the same more modest claim: a research record produced under P1âP8 is more auditable, more contestable, and harder to misuse as a compliance proxy than one produced under current norms, regardless of whether it achieves certainty. These asymmetries mean that alignment researchers cannot simply import the evaluaHon norms of mainstream AI. Instead, they must elevate standards of evidence and rigor to account for the higher stakes and trickier epistemics. We next analyse specific dimensions along which mainstream AI research pracHces versus alignment needs diverge and idenHfy gaps that any adequate epistemic framework for AI safety must address. 3. Gap Analysis: Mainstream Norms vs. Alignment Needs The two asymmetries of §2 describe what alignment must demonstrate. They do not by themselves specify where mainstream research pracHce falls short. To go from the asymmetries to diagnosHc leverage, we introduce five cross-cuOng dimensions along which mainstream norms and alignment needs diverge. Each dimension is an instrument that may surface failure modes on either asymmetry: the same gap (for example, weak verificaHon) can allow a capability-profile failure to go undetected and allow a risk-profile failure to go unquanHfied. The five dimensions are therefore not further instances of the two asymmetries but the observable channels through which both asymmetries are violated in pracHce. Mainstream AI research is built around average-case opHmisaHon (benchmark scores, mean accuracy, expected reward) and is typically accompanied by selecHve post-hoc artefact sharing, limited independent verificaHon, shallow uncertainty reporHng, and weak incenHves to reward negaHve results or adversarial criHque (Russakovsky et al. 2015; Wang et al. 2019; Gundersen and Kjensmo 2018; Henderson et al. 2018; Raff 2019; Pineau et al. 2021; MunafĂČ et al. 2017). Alignment has to invert these prioriHes: opHmise for worst-case (tail) safety, demand radical up-front openness, invite adversarial mulH-team replicaHon and stress-tesHng, embrace mulH-layered uncertainty accounHng, and foster insHtuHonalised scepHcism (Bostrom 2014; Shevlane et al. 2023; Hubinger et al. 2019; Amodei et al. 2016; Madry et al. 2018; Leveson 2012). The five dimensions below are the places where this inversion either happens or fails to happen. 1. What is being op.mised against (failure-mode targeHng vs. mean-performance targeHng) 2. Visibility of evidence (transparency) 3. Trustworthiness of evidence (verificaHon & replicaHon) 4. Treatment of unknowns (uncertainty accounHng) 5. Enforcement culture (incenHves and norms for rigour). Each dimension corresponds to well-documented failure modes in other high-reliability fields. Mapping these gaps yields a focused rubric for diagnosing where convenHonal AI research pracHces break down when transplanted into the alignment domain. 6 3.1 Bibliometric baseline (preregistered protocol) The five gap dimensions above are derived from a structured comparaHve synthesis of safety- engineering literature, meta-scienHfic work on reproducibility and selecHve reporHng, and published criHques of AI evaluaHon pracHce (Henderson et al. 2018; Pineau et al. 2021; MunafĂČ et al. 2017; Reuel et al. 2024; Bean et al. 2025; Wan et al. 2025). A reasonable addiHonal quesHon, beyond the synthesis itself, is whether these failure modes are also detectable as observable prevalence rates in the published alignment literature. To address this directly, we have registered a preregistered bibliometric baseline study on the Open Science Framework (10.17605/OSF.IO/45YR6). The protocol specifies a straHfied random sample of n â„ 30 alignment papers across four sub-areas (RLHF; red- teaming and safety evaluaHon; mechanisHc interpretability; robustness), five binary observable markers corresponding to E1âE5, a bounded-evidence coding rule, and double coding with reliability reporHng. Wilson 95% confidence intervals are reported on prevalence rates per marker and per stratum. Pilot fragment status: The single-coded pilot reported below is a parHal implementaHon against the registered protocol, not a subsHtute for it. We are conducHng a pilot fragment of n = 10 papers (4 RLHF, 3 red-teaming/safety evaluaHon, 2 mechanisHc interpretability, 1 robustness) during the manuscript revision window, single-coded by the principal invesHgator using the registered codebook. The pilot fragment is reported here to give an indicaHve reading of marker prevalence; it is not a subsHtute for the full study. The full sample (n â„ 30, double-coded with one externally- recruited second coder) is in progress and will be reported in a follow-up paper. Pilot codings are retained and incorporated into the full sample; deviaHons from the registered protocol, if any, will be logged on OSF. Indica.ve pilot results. Coding of the n = 10 pilot fragment yielded the following marker prevalence rates (proporHon of papers scoring 1, with Wilson 95% confidence intervals): E1 worst-case or adversarial reporHng, 40.0% (95% CI 16.8â68.7%); E2 transparency and disclosure, 10.0% (95% CI 1.8â40.4%); E3 independent verificaHon, 0.0% (95% CI 0.0â27.8%); E4 explicit residual risk or uncertainty, 90.0% (95% CI 59.6â98.2%); E5 reporHng completeness, 10.0% (95% CI 1.8â40.4%). The pilot's strongest signal, and the most informaHve result for the gap analysis, concerns E3: independent verificaHon was absent in every paper coded. None of the ten papers cites an independent replicaHon, a third-party audit, or an external red-team study of its central claim. Even allowing for the wide confidence interval that necessarily accompanies a single-coded n = 10 pilot, the upper Wilson bound of 27.8% is sufficient to establish that the pracHce cannot be common in alignment research, even under favourable assumpHons about under-counHng. This is consistent with the gap-analysis predicHon that insHtuHonalised independent verificaHon is largely absent from the field, and it is the empirical observaHon on which the case for P4 (Independent VerificaHon) most directly rests. Second, E4 is observed at 90%, the highest prevalence among the five markers: substanHve residual-risk or limitaHons discussion appears almost universally. We surface this finding rather than pass over it: the dimension on which alignment papers most consistently score well is the open statement of limitaHons, and that should be acknowledged. The qualifier needed is that E4 is a deliberately low bar (i.e., presence of a substanHve limitaHons secHon beyond boilerplate) and meeHng it does not imply that the field handles uncertainty well in any deeper sense. The retrospecHve rubric audit reported in §7 finds that P7 (Reflexive Uncertainty AccounHng), which assesses structured uncertainty registers and quanHfied residual-risk statements, scores near zero across all audited sub-areas. The two readings are coherent: alignment papers acknowledge limitaHons in prose but rarely provide the structured uncertainty artefacts P7 specifies. Third, the sub-community paperns are intelligible and corroborate a finding from the rubric audit. E1 prevalence is 100% in the red-teaming and robustness strata (where adversarial tesHng is itself the 7 paper's topic) and 0% across the RLHF and mechanisHc-interpretability strata, indicaHng that worst- case evaluaHon in alignment research is concentrated in a small set of sub-communiHes rather than diffused as a general pracHce. The §7 audit independently observed the same sub-community- specific paperning across its own sample, with red-teaming papers scoring higher on P6 by design and artefact-release pracHces varying systemaHcally across sub-areas. The two pieces of empirical work, conducted on different samples and with different instruments, agree on this point, which strengthens both. Three caveats apply to the pilot itself: (i) n = 10 is small and the confidence intervals are correspondingly wide, parHcularly for E1 (16.8â68.7%); (i) the pilot is single-coded and reliability cannot be assessed at this stage; (i) the markers are conservaHve observable proxies, not direct measurements of underlying epistemic pracHce. We treat these figures as indicaHve of the paperns the full study will resolve more precisely, not as field-wide claims; the full study (n â„ 30, double-coded) is in progress. Even seOng the pilot aside, the dimensions idenHfied in this secHon warrant apenHon because they correspond to documented failure modes in adjacent high-reliability fields, not merely because we have postulated them. Mapping these gaps yields a focused rubric for diagnosing where convenHonal AI research pracHces break down when transplanted into the alignment domain. 3.2 What is being op6mised against Mainstream AI opHmizes average-case performance, higher accuracy, lower error, beper reward, while alignment must minimize worst-case (tail) risk and furnish high-confidence evidence that no catastrophic failure will occur. Standard evaluaHons privilege means improvements and rarely incenHvize exhausHve searches for rare failure modes; a system can post excellent averages while concealing a handful of extreme errors. Alignment researchers, by contrast, must acHvely hunt those one-in-a-thousand (or rarer) failures and design them out. The asymmetry mirrors high-reliability engineering, where disasters have followed the normalizaHon of deviance that ignored or downplayed anomalies exemplified by the Challenger launch decision (Vaughan 1996; PresidenHal Commission 1986; Feynman 1986; Perrow 1999; Turner 1976).. Accordingly, benchmark leaderboards must give way to systemaHc stress-tests that drive models into extreme regimes, coupled with adversarial red-team campaigns that deliberately hunt for edge-case failures. In this safety-first mindset, any non-zero catastrophic failure rate is presumpHvely unacceptable unHl exhausHve analysis shows otherwise. The relevant high-reliability precedent is aerospace design-assurance pracHce, which treats catastrophic-failure probabiliHes below 10â»âč per flight-hour as a target for the assurance argument rather than a measured achievement (FAA AC 25.1309-1B). Alignment cannot compute against an analogous number: it lacks base-rate data, staHonary system dynamics, and a tractable enumeraHon of failure modes. What it can inherit from aerospace is the underlying principle of evidence-proporHonality: the stringency of the evidence required for a safety claim should scale with the severity of the outcomes the claim rules out. Yoshua Bengio captures the resulHng posture: demand âvery strong evidence before concluding there is nothing to worry aboutâ (Bengio 2024), and invoke the precauHonary principle given âthe extreme severity and unknown likelihood of catastrophic risksâ (Bengio 2025). Shevlane et al. (2023) formalise this stance, proposing evaluaHon frameworks that centre tail-risk assurance rather than mean-case performance. 8 3.3 Transparency Mainstream AI onen withholds code and data unHl aner publicaHon, blocking independent audits and leOng errors lurk unseen. Alignment research cannot accept that risk. In safety-criHcal domains, opacity is itself a hazard, and safety standards, from IEC 61508 to DO-178C, require fully reviewable evidence (Leveson 2012; Bloomfield and Bishop 2010). Alignment work should therefore pracHse extreme transparency: preregister protocols, release datasets, checkpoints, and evaluaHon pipelines at the outset, and host them in FAIR-compliant repositories for conHnuous peer audiHng. Academic pracHce is already partway there: preprint servers and arHfact tracks make code and data available earlier and more systemaHcally than before, and the most concerning cases are increasingly concentrated in the commercial tail, where the most capable systems sit behind confidenHality walls. Precedents for the standard we propose include clinical-trial registries, genomics pre-publicaHon data sharing, and ecology open-data mandates, together with frameworks such as the TOP Guidelines, FAIR principles, and reproducibility manifestos that define the minimum bar (De Angelis et al. 2004; Wilkinson et al. 2016; MunafĂČ et al. 2017). Only through such real-Hme, community-level scruHny can alignment claims achieve the ultra-low residual-risk assurance that high-integrity systems demand. 3.4 Verifica6on Mainstream AI typically treats single-team experiments, rouHne cross-validaHon, and standard peer review as adequate evidence; independent replicaHons are uncommon and incenHves to break published results are weak, so findings that clear cursory review onen ossify into the literature (Gundersen and Kjensmo 2018; Henderson et al. 2018; Raff 2019; Pineau et al. 2021; Langley 2019). By contrast, alignment research must regard every unchallenged safety claim as a potenHal single point of failure: adversarial replicaHon, coordinated mulH-team stress-tesHng, and structured red- teaming should be the default, with third-party audits and independent verificaHon bodies rewarded at least as highly as original discovery (Shevlane et al. 2023; Hubinger et al. 2019; Amodei et al. 2016; Madry et al. 2018; Bloomfield and Bishop 2010; DARPA 2019). Because a false sense of security could be catastrophic in this domain, results should remain strictly provisional unHl they have survived such hosHle scruHny (Bostrom 2014; Carlsmith 2022). 3.5 Uncertainty Accoun6ng Mainstream ML papers rouHnely report confidence intervals or p values yet seldom confront epistemic uncertainty such as unknown unknowns, briple modelling assumpHons, and distribuHon shin, resulHng in overconfident claims drawn from a single dataset or narrow regime (Gundersen & Kjensmo 2018; Pineau et al. 2021; NaHonal Academies of Sciences, Engineering, and Medicine 2019). Robustness to new seOngs is rarely stress tested, so findings that clear cursory peer review onen ossify even though they hold only under restricted condiHons (Ioannidis 2005; Langley 2019). For alignment work this complacency is dangerous: with scant data on advanced AI behaviour, ignorance itself becomes a key variable (Bostrom 2014; Shevlane et al. 2023). Alignment studies should therefore pracHse mulH-layered uncertainty accounHng that (i) quanHfies staHsHcal, model, and structural uncertainty using Bayesian or ensemble methods to yield credible intervals; (i) probes worst-case tails via scenario analysis; and (i) runs periodic calibraHon audits to test forecast accuracy (HoeHng et al. 1999; GneiHng and Ranery 2007; Lempert et al. 2003; Tetlock & Gardner 2015). Where ground truth is sparse or non-existent, structured expert judgment (SEJ) offers a calibrated, auditable procedure for eliciHng and aggregaHng probabiliHes with performance-based weights (Cooke 1991). Climate scienceâs mulH-model ensembles provide a mature precedent for integraHng structural uncertainty across modelling choices (KnuO et al. 2010). Each result should carry a blunt limits statement, e.g., âThis holds under A, B, C; beyond that, uncertainty balloonsâ, and an epistemic audit trail documenHng every modelling choice (Mitchell et al. 2019; Gebru et al. 2021). 9 Crucially, researchers must quanHfy residual risk that plausible failures remain undetected, guarding against the fallacy that absence of evidence is evidence of absence (Cohen 1988; Altman & Bland 1995). 3.6 Enforcement culture: organised scep6cism Mainstream AI culture leans on informal, âlight-touchâ peer review; negaHve results and sharp criHques are nominally welcomed but weakly rewarded, breeding polite deference and group-think (Langley 2019; Pineau et al. 2021; MunafĂČ et al. 2017). Such environments normalise deviance where anomalies are waved away so long as nothing visibly breaks, a dynamic that helped precipitate both the Challenger and Fukushima failures (Vaughan 1996; NaHonal Diet of Japan 2012). More broadly, ânormal accidentsâ theory underscores how Hghtly-coupled, high-complexity systems tend toward unanHcipated failure interacHons, reinforcing the need for insHtuHonalised scepHcism (Perrow 1999). Turner âs classic analysis of the âincubaHonâ of disasters, where latent errors accumulate under organisaHonal blind spots, likewise moHvates formal adversarial review before claims are accepted (Turner 1976). Alignment work, carrying civilisaHon-scale stakes, must instead insHtuHonalise scepHcism. Major safety claims should automaHcally trigger adversarial red-team reviews, independent groups tasked with assuming the result is wrong and proving it (Shevlane et al. 2023; Hubinger et al. 2019). Public criHque channels such as predicHon markets, âbreak-itâ contests, replicaHon bounHes, can crowd-source expert doubt and reward flaw-finders (Hanson 1995; Camerer et al. 2018). Mirroring high-reliability engineering, every criHcal conclusion should face fresh-eyes stress tests and line-by-line scruHny before it is accepted as knowledge (Leveson 2012; ISO/IEC 15026-2:2011). Table 1 summarizes these five gaps, highlighHng how current norms are inadequate and poinHng to the new pracHces needed in alignment (many inspired by other disciplines). Table 1. Gap analysis: where mainstream AI research norms fall short of alignment requirements Dimension Mainstream AI norm (status quo) Resul5ng vulnerability / âgapâ Alignment specific norm or prac5ce required Cross domain precedent Op-misa-on target (performance vs. safety) Average-case metrics (accuracy, reward, benchmark rank) dominate evalua-on. Rare but catastrophic errors go undetected and un-mi-gated. Op-mise for tail-risk minimisa,on; treat any non-zero catastrophic failure probability as unacceptable un-l exhaus-vely analysed and reduced. Aerospace design- assurance target of †10â»âč catastrophic failures per flight-hour (used as mo-va-ng analogy, not as computable target for alignment). Evidence transparency Code, data, and protocols released late, par-ally, or behind NDAs. Hidden assump-ons and corner case bugs escape external audit. Extreme transparency: preregistered protocols, immediate release of datasets, checkpoints, and evalua-on pipelines in FAIR repositories. Clinical trial preregistra-on and mandatory results repor-ng; TOP Guidelines level 3 data sharing. Verifica-on & replica-on Single team experiments plus light peer review; lible incen-ve for replica-on or red teaming. Brible results ossify; false security builds around unchallenged findings. Mandated adversarial mul- team replica-on and red team âbreak itâ exercises before accep-ng any safety claim. DARPA SCORE programme; cybersecurity penetra-on tes-ng. Uncertainty accoun-ng Error bars on convenience metrics; epistemic uncertainty (model, data, shih) rarely quan-fied. Overconfident conclusions hold only under narrow condi-ons; âunknown unknownsâ ignored. Mul- layer uncertainty quan-fica-on: Bayesian/ensemble credibility intervals, scenario stress tests, con-nual calibra-on Mul- model climate projec-ons; intelligence analysis calibra-on labs. 10 audits, explicit limits statements. Culture & incen-ves (organised scep-cism) Informal âpoliteâ peer review; nega-ve results and strong cri-ques undervalued. Groupthink, normalisa-on of deviance, missed warning signs (Challenger type failures). Ins-tu-onalise scep-cism: automa-c red team reviews, predic-on markets, replica-on boun-es; career credit for flaw finding and replica-on. Independent nuclear safety verifica-on teams; pharmaceu-cal challenge studies and post market surveillance. 4. Bridging the Gaps ExisHng governance regimes including NIST and ISO standards, the EU AI Act, open-science reforms, provide valuable pieces of the puzzle, yet none delivers a complete epistemic framework for alignmentâs prevenHve, tail-risk mission. They emphasise organisaHonal process over experiment-level worst-case tesHng, stop short of enforcing reproducibility or staHsHcal power during early research, and omit tools such as systemaHc red-teaming or formal uncertainty audits that are indispensable when an adapHve model might conceal its own faults (Shevlane et al., 2023; Amodei et al., 2016). Closing this gap requires drawing on four enabling tradiHons, each covering a disHnct slice of the five diagnosHc dimensions introduced above (Table 2). We keep philosophy of science and quanHtaHve risk analysis separate because they do different work: the former specifies when a claim is well- supported (severe tesHng, falsificaHon), whereas the laper provides tools for represenHng and quanHfying uncertainty (scenario analysis, Bayesian model averaging, structured expert judgment). Combining them obscures that ECAISA synthesises a philosophical stance with a set of engineering methods rather than a single tradiHon. Table 2. Four enabling tradiAons and the epistemic gap dimensions each most directly addresses. Open science, safety engineering, philosophy of science (severe tesAng), and quanAtaAve risk analysis each contribute disAnct methodological resources to ECAISA. No single tradiAon spans all five dimensions; the synthesis is the contribuAon. Tradi5on Core contribu5on Dimensions most directly addressed Open Science: (radical transparency & exhaus-ve reproducibility) Preregistra-on, FAIR-compliant releases, ar-fact review, and large-scale replica-on unlock con-nuous external scru-ny. (i) Visibility of evidence; (i) Trustworthiness via reproducibility; (v) Enforcement culture through credit for sharing and replica-on. Safety Engineering: (precau-on, rigorous V&V, worst-case design) Defence-in-depth, independent audits, and safety cases shih the objec-ve from mean performance to near-zero catastrophic risk. (i) Op,misa,on target,tail-risk minimisa-on; (i) Trustworthiness via formal V&V; (v) Enforcement culture through mandatory checks and cer-fica-on. Philosophy of Science: (severe tes-ng and falsifica-on) Severe tes-ng (Popper, Mayo) specifies when a claim is well-supported: by surviving genuine abempts to refute it rather than by accumula-ng confirma-ons. (i) Trustworthiness of evidence via adversarial tes-ng; (v) Enforcement culture by ins-tu-onalising doubt (adversarial reviewers, predic-on markets). Quan5ta5ve Risk Analysis: (uncertainty representa-on and scenario analysis) Bayesian and ensemble modelling, scenario analysis, structured expert judgment, and calibra-on audits (Cooke 1991; Hoe-ng et al. 1999; Lempert et al. 2003) turn ignorance itself into a representable and auditable object. (iv) Treatment of unknowns; (i) Op-misa-on target via quan-fied tail-risk assessment. No single tradiHon spans all five dimensions, but their synthesis can. Open science maximises transparency but lacks worst-case engineering targets; safety engineering delivers assurance rigour but not radical openness; severe tesHng specifies a discipline of claim evaluaHon but does not prescribe disclosure protocols; quanHtaHve risk analysis represents uncertainty but does not specify 11 verificaHon thresholds. Alignment therefore needs an integrated epistemic code that treats verifiability, adversarial robustness, and uncertainty quanHficaHon as non-negoHable safety requirements. In the next secHon we operaHonalise this synthesis by proposing eight Epistemic Principles, collecHvely the AI Alignment Epistemic Code, that embed the transparency of open science, the assurance discipline of safety engineering, the scepHcal rigour of severe tesHng in philosophy of science, and the uncertainty-representaHon tools of quanHtaHve risk analysis. For each principle we specify concrete obligaHons and show how it directly remediates the gaps idenHfied across the five dimensions. 5. Proposed Epistemic Principles Building on the gap analysis and the four enabling tradiHons (open science, safety engineering, philosophy of science, and quanHtaHve risk analysis), we now translate those ideas into ECAISA, eight concrete non-opHonal duHes for any research that claims to advance AI alignment. Each principle defines an obliga.on (what must be done) and a ra.onale (why it mapers for tail-risk safety), see Table 3. Table 3: Proposed Epistemic Principles ID Obliga5on (what to do) Ra5onale (why it maMers) P1. Radical Transparency Preregister hypotheses and release code, data, model checkpoints, and evalua-on pipelines in FAIR repositories before publica-on (subject only to clearly documented info hazard excep-ons). Con-nuous external audit is the only reliable way to surface hidden failure modes in systems that may act decep-vely. P2. Comprehensive Integrity Report all results, including null/nega-ve findings, and disclose funding and conflicts; secure tamper evident logs for the full research lifecycle. Suppressing inconvenient data distorts the risk picture; integrity lapses in safety science can cost lives. P3. Evidence Propor5onal Claims Match the strength of every claim to quan-fied support (effect sizes, credible intervals, or formal proofs); preregister analysis plans to block post hoc spin. Overclaiming breeds false confidence: in tail risk senngs, absence of evidence is not evidence of safety. P4. Reproducibility + Adversari al Verifica5on Provide full replica-on packages and commission at least one independent, hos-le replica-on for any major safety claim. Robustness comes from surviving abempts to break the result, not from a single teamâs success case. P5. FAIR Stewardship of Ar5facts Treat datasets, benchmarks, and trained models as community assets: assign DOIs, supply rich metadata, and maintain versioned releases. Well curated ar-facts enable follow on audits, compara-ve studies, and faster collec-ve progress. P6. Ins5tu5onalised Scep5cism Make red team review the default: publish cri-que reports or âbreak itâ results alongside the primary paper; reward flaw finders. Organised doubt prevents group think and the normalisa-on of deviance that doomed past high risk projects. P7. Reflexive Uncertainty Accoun5ng Maintain a living uncertainty register that quan-fies sta-s-cal and epistemic unknowns, updates with new evidence, and reports calibra-on metrics. Explicit ignorance tracking guards against the illusion of safety when evidence is sparse and models are adap-ve. P8. Epistemic Hygiene Foster rou-nes that surface bias and confla-on of specula-on with fact, devilâs advocate roles, cross disciplinary audits, clear labelling of conjecture. A clean knowledge ecosystem is a prerequisite for reliable cumula-ve science, especially under existen-al stakes. A pracHcal quesHon arises: when should a venue, funder, or laboratory consider a given ECAISA principle to be fulfilled? To address this we specify a three-level ordinal rubric (Table 4) that can be applied to a paper or research artefact on the basis of what a reader or reviewer can verify from the 12 paper text, its appendices, and author-linked artefacts. A score of 2 (âmeetsâ) is assigned only when the principleâs operaHonal threshold is explicitly evidenced; a score of 1 (âparHalâ) indicates that the pracHce is present but falls short of the threshold; a score of 0 (âabsentâ) indicates no substanHve evidence of the pracHce beyond generic prose. When deciding between 1 and 2, prefer 1 unless the âmeetsâ criterion is explicitly evidenced; when deciding between 0 and 1, prefer 0 unless some substanHve evidence is present. The full codebook with per-principle coding rules, He-break rules, annotaHon-sheet fields, and the inter-rater reliability computaHon is given in Appendix A. Table 4. Summary scoring rubric for ECAISA principles. The full codebook is in Appendix A. Principle Threshold for a âmeetsâ score (2) P1 â Transparency Sufficient artefacts released (code, data, evalua:on pipeline) to enable meaningful scru:ny, or a concrete disclosure protocol under the ladder (Table 5) with: (i) explicit list of withheld items, (i) jus:fica:on, (i) mi:ga:ons and access pathway, and (iv) a re-evalua:on condi:on. P2 â Full repor:ng, COI, provenance P2a: failures and nulls reported alongside successes, with pre-specified primary outcomes. P2b: all funding and relevant conflicts named, with access asymmetries an independent replicator would not share disclosed. P2c: opera:onal correc:on infrastructure (public issue tracker, errata policy), and for T3 claims (see Table 6), tamper-evident provenance for core lifecycle events. P3 â Evidence-propor:onal claims Measurable success criteria with: clear metrics and baselines aligned to claims; sufficient experimental detail; and at least one robustness- oriented element propor:onate to the claim. P4 â Independent verifica:on Central claim subjected to at least one of: independent replica:on by a separate team, structured third-party red-teaming with reportable traces, or a release package designed for adversarial verifica:on with a clear protocol and evidence of uptake. P5 â Ar:fact stewardship Artefacts stewarded as citable, durable research objects: versioned release, stable iden:fier (e.g. DOI), environment specifica:on, licence and provenance notes, and a maintenance/ownership statement. P6 â Red-teaming Structured adversarial evalua:on with: explicit threat model and aYack surface; systema:c test design and constraints; clear repor:ng of successes and failures; and reproducible artefacts or documented excep:ons. P7 â Uncertainty accoun:ng Explicit uncertainty artefacts including a structured uncertainty register (assump:ons, known unknowns, failure modes), a residual-risk statement with plausible falsifiers, and calibra:on or sensi:vity analyses where applicable. P8 â Decision relevance Claims explicitly situated in an ins:tu:onal decision context: clear claim boundaries (what this does and does not imply), explicit COI/funding and access-asymmetry disclosure, and concrete decision-facing guidance with reliance condi:ons and update triggers. 13 5.1 How the ECAISA code closes the five epistemic gaps Figure 1. Mapping the eight ECAISA principles onto the five epistemic dimensions. Figure 1 (and the accompanying table) briefly indicates how the eight ECAISA principles distribute themselves across the five gaps idenHfied in SecHon 3. It also indicates how the eight ECAISA principles collecHvely seal the epistemic gaps idenHfied earlier. Each principle occupies a disHnct papern of green dots, showing that while none spans every column, every column is nevertheless covered several Hmes over. This deliberate overlap mapers: if radical transparency (P1) is delayed or parHal in a given project, the same evidence sHll falls under FAIR stewardship requirements (P5) and remains open to red-team disclosure under insHtuHonalised scepHcism (P6). Redundancy ensures that no single lapse in pracHce can reopen a gap. One principle, Evidence-ProporHonal Claims (P3), plays a special integraHve role. It is the only pillar that engages four of the five columns: opHmisaHon target, transparency, verificaHon, and uncertainty. By binding claim strength to quanHfied support it Hes the tail-risk objecHve directly to the mechanisms that make evidence trustworthy and uncertainHes explicit, thereby turning âextraordinary claims require extraordinary evidenceâ into an operaHonal rule rather than a slogan. Read across the grid and the message is unambiguous: every vulnerability idenHfied in Figure 1 is countered by at least one, and usually several, mandatory duHes. ECAISA therefore supplies 14 complete dimensional coverage, giving laboratories, journals, and funders a coherent standard that matches the tail-risk profile and prevenHve mission of alignment research. 5.2 Disclosure and 6ered applicability ECAISA does not require that every artefact be made fully public. Two concerns weigh against unqualified openness: informaHon-hazard risk (releasing material that could meaningfully increase misuse capability) and legiHmate commercial confidenHality (protecHng trade secrets and compeHHve posiHon). We address both through a single mechanism: a four-level disclosure ladder (Table 5) that decouples evidence transparency from capability transparency. A study can saHsfy the disclosure principle P1 at level L2 by sharing weights, prompts, or evaluaHon traces with veped replicators under responsible-disclosure terms without releasing them publicly, and at L3 by registering the existence and raHonale of a restricHon now and releasing artefacts once risks reduce or miHgaHons exist. The pharmaceuHcal analogy is instrucHve: drug companies do not publish synthesis routes but do publish trial protocols, adverse-event data, and staHsHcal pre-analysis plans, and it is these laper artefacts, the evidenHary record, rather than the implementaHon, that ECAISA targets for alignment. Table 5. Disclosure ladder for implemenAng P1 while managing informaAon hazards and commercial confidenAality. Level What is released L0 â Public Code, data (as legally permissible), evalua:on pipeline, aggregate results, uncertainty-register summary. L1 â Public with redac:ons Public artefacts with capability-relevant prompts, weights, or aYack details redacted; release ra:onale logged. L2 â Controlled access Sensi:ve artefacts shared with veYed third par:es (replicators or red- teamers) under responsible-disclosure terms. L3 â Delayed release Time-locked release; publish existence and ra:onale now, release artefacts a_er risks reduce or mi:ga:ons exist. Not every alignment claim warrants the same evidenHary burden. A Hered applicability scheme (Table 6) scales ECAISA obligaHons with the potenHal real-world impact of the claim: exploratory and rouHne research claims (T0âT1) can saHsfy minimum-viable pracHces, reserving the full P1âP8 protocol for claims that will inform deployment or governance decisions (T3). This preserves the stringency of ECAISA where it mapers most while avoiding disproporHonate burden on early-stage or hypothesis-generaHng work and reduces the perceived cost of adopHon for industry labs doing exploratory research. Table 6. Tiered applicability of ECAISA: scope of obligaAons by claim type and stakes. Tier When applicable Minimum ECAISA expectaAons T0 â Exploratory Early conceptual work; hypothesis genera:on; non-deployment claims. Label conjectures clearly; maintain an uncertainty register (P7-lite); disclose conflicts of interest (P2b-lite). T1 â Research claims Empirical alignment claims used to steer research agendas. Preregistra:on or analysis plan (P1, P3); full repor:ng including nulls (P2a); replica:on package (P4). T2 â Safety claims Claims that a method reduces catastrophic or decep:ve failure risk. Adversarial verifica:on or independent replica:on (P4); red-team report (P6); disclosure ladder (P1); uncertainty audit (P7). T3 â Deployment- cri:cal Claims used to jus:fy deployment or governance decisions. Full P1âP8 unless a documented infohazard excep:on applies; auditable artefact stewardship (P5); governance and an:-gaming checks. 15 5.3 Infohazard adjudica6on Principle P1 and the disclosure ladder together specify that some artefacts may legiHmately be withheld, but do not specify who decides, under what criteria, or with what accountability. Without such a procedure, the âdocumented infohazard excepHonâ risks becoming a cover for selecHve opacity. The following protocol operaHonalises the adjudicaHon process; it is intended as a minimum floor, not a ceiling. Decision body. A designated infohazard review commipee is convened for each restricHon decision. The minimum composiHon is the principal invesHgator, an independent safety reviewer unaffiliated with the study, and at least one external domain expert. For organisaHonal research (industry labs, research insHtutes), the internal safety team subsHtutes for the external expert where domain knowledge is proprietary, but a named external reviewer is sHll required for T2 claims and above. For T3 claims, a formal body analogous to an IRB is required, with standing members and recorded minutes. Decision criteria. The commipee completes and signs a five-part template for each restricHon: (i) threat model (who could misuse the artefact, how, and at what scale); (i) plausibility and expected harm magnitude; (i) miHgaHon feasibility (redacHon, aggregaHon, controlled access under L2); (iv) the disclosure level applied under the ladder; and (v) a re-evaluaHon date or condiHon aner which broader release is considered. The five elements are required; a missing element is itself a reason to withhold approval of the restricHon. Documenta.on. A public log entry records that a restricHon has been applied, with minimal metadata: the disclosure level applied, the re-evaluaHon date, the categories of artefact withheld, and the signing members of the commipee. The substanHve restricted content remains withheld. Publishing the existence and basic raHonale of a restricHon is sufficient to make the decision auditable without defeaHng the restricHon itself. Appeal and contestability. A restricted artefact can be challenged through the public correcHon channel (see anH-gaming mechanisms, §7.1). ParHes may contest either the restricHon itself (as over- broad or unsupported by the threat model), the commipee composiHon (as insufficiently independent), or the conHnuaHon of the restricHon past its re-evaluaHon date. A contested restricHon triggers a second review by a fresh commipee at least half of whose members were not party to the original decision. Accountability. Commipee members are named in the publicaHon and sign the public log entry. This mirrors the insHtuHonal review board apribuHon model: the decision is apributable to specific individuals, not to the insHtuHon as an anonymous actor. Naming creates a reputaHonal surface that incenHvises the commipee to reason carefully and document thoroughly; it also gives a challenging party a concrete counterparty. For T3 claims, the commipee chair signs an explicit apestaHon that the five decision criteria were completed in good faith. This protocol does not purport to solve every dual-use dilemma in alignment research. What it does is ensure that when transparency and safety conflict, the resoluHon is a documented and contestable decision rather than an informal discreHon. The infohazard excepHon becomes a governance artefact in its own right, with the same kind of auditability ECAISA asks of the rest of the research record. 5.4 Implementa6on Outlook We envisage journals, funders, and labs adopHng ECAISA as a visible badge of rigor, analogous to clinical-trial preregistraHon numbers or DO-178C designaHons in avionics. Principles P1âP5 can be operaHonalised immediately by extending exisHng open-science and V&V infrastructure; P6âP8 require cultural commitment but draw on well-tested mechanisms (penetraHon-test contracts, predicHon markets, bias-awareness training). 16 6. Related Work: Posi0oning the Proposed Epistemic Code Our proposal for stricter epistemic norms in AI safety draws on, but also goes beyond, scholarship in AI alignment, philosophy of science, sociology of research, metascience, and emerging governance frameworks. In what follows we posiHon ECAISA within this landscape, showing where exisHng literature already supports elements of our code, where it remains silent, and how our principles close those gaps. By tracing contribuHons from canonical academic sources and influenHal grey literature alike, we clarify both the intellectual lineage of our approach and the disHncHve advances it offers to the conversaHon on trustworthy, high-stakes AI research. 6.1 Dis6nc6ve contribu6on and governance crosswalk A natural reviewer quesHon at this point is whether ECAISA amounts to more than a bundling of pracHces already required, in some form, by exisHng frameworks such as the TOP Guidelines, the FAIR principles, NIST AI RMF 1.0, ISO/IEC 42001, ISO/IEC 15026-2, and the EU GPAI Code of PracHce. Our posiHon is that ECAISAâs contribuHon is not in invenHng individual pracHces but in specifying the minimal claim-level bundle that collecHvely closes the five epistemic gaps, together with anH-gaming infrastructure that none of the cited frameworks imposes. Four points make this concrete. Research-phase scope. ExisHng governance frameworks (NIST AI RMF, ISO/IEC 42001, EU AI Act and its GPAI Code of PracHce), evaluaHon reporHng standards for dangerous-capability evaluaHons (STREAM, McCaslin et al. 2025) and benchmark assessment frameworks (BeperBench, Reuel et al. 2024), transparency metrics (FMTI, Wan et al. 2025), safety-case methodology (Buhl et al. 2024; Hilton et al. 2025; Goemans et al. 2024; Barrep et al. 2025), and audiHng proposals (Raji et al. 2020) operate at organisaHonal, deployment-assurance, or model-report levels. ECAISA occupies the upstream layer: the epistemic pracHces through which alignment claims are first produced and recorded, before they feed any safety case or transparency index. The relaHonship is sequenHal rather than compeHHve: ECAISA-compliant research artefacts would saHsfy many STREAM disclosure criteria for safety-relevant benchmark evaluaHons as a downstream consequence and would contribute to higher FMTI scores, but ECAISA's requirements extend further upstream into preregistraHon, independent adversarial verificaHon, and structured uncertainty registers that neither STREAM nor FMTI addresses. Integra.on of tradi.ons. We combine preregistraHon and FAIR-compliant release (open-science verifiability) with worst-case targets, defence in depth, and formal safety-assurance methodology (safety-engineering conservaHsm) in a single code. ExisHng open-science checklists lack worst-case targets; safety-engineering frameworks such as ISO/IEC 15026-2 and DO-178C lack radical disclosure obligaHons. ECAISA requires both simultaneously. The synthesis is the contribuHon: P1âP5 together specify what neither tradiHon specifies alone. Auditability rather than cer.fica.on. Rather than apempHng to cerHfy system safety, which is not achievable for open-ended systems (Narayanan and Kapoor 2023), ECAISA specifies the condiHons under which a safety claim can be independently reconstructed, contested, and updated. This is a different governance target from cerHficaHon: auditability is weaker (it does not guarantee the claim is correct) but more robust (it does not depend on closing the denominator of the failure probability). No exisHng governance framework explicitly commits to auditability rather than cerHficaHon as its primary target. An.-gaming infrastructure. Any compliance regime is vulnerable to ritualisaHon (see §7.1). ECAISA pairs its obligaHons with an anH-gaming bundle â spot audits, lopery-based independent verificaHon, hash manifests, public correcHon channels, logged disclosure excepHons, red-team minimum fields, and access-asymmetry disclosure â that none of NIST, ISO, or the EU GPAI Code 17 currently specifies. Without this infrastructure, even well-wripen obligaHons can collapse into checklist-filling. Table 7 makes the relaHonship explicit at the principle level: each of the eight ECAISA principles is mapped to the nearest hook in NIST AI RMF 1.0, ISO/IEC 42001, and the EU GPAI Code of PracHce, and the rightmost column states the disHncHve obligaHon ECAISA adds beyond what each of those frameworks currently specifies. The crosswalk is intended as an integraHon aid, not a compliance claim: the hooks indicate conceptual alignment, not formal conformity, and organisaHons should conHnue to treat the underlying standards as authoritaHve for their own purposes. Read verHcally down the rightmost column, the disHncHve ECAISA obligaHons are not a relabelling of exisHng duHes. Compliance with NIST AI RMF or ISO/IEC 42001 does not require preregistraHon of analysis plans, published adversarial verificaHon, a structured uncertainty register, access-asymmetry disclosure, tamper-evident provenance logs, or a commipee-adjudicated infohazard excepHon. Each of these is a non-trivial addiHon, and each addresses a specific failure mode of the research record that the organisaHonal frameworks are not designed to catch. Stated differently: an organisaHon could be fully compliant with NIST, ISO, and the EU GPAI Code and sHll produce research that ECAISA would score poorly. 18 Table 7. Governance crosswalk. Hooks indicate conceptual alignment; ECAISA does not itself establish compliance with any legal or standards-based regime. ECAISA principle NIST AI RMF 1.0 ISO/IEC 42001 EU GPAI Code of Prac5ce What ECAISA adds P1 â Transparency & disclosure Map 1.1 (context); Govern 1.2 (documenta-on) §6.1 Documented informa-on; §8.2 AI risk assessment Art. 11 Technical documenta-on; Art. 53 Transparency Preregistra-on of protocols; FAIR-compliant release at study outset; disclosure ladder with documented infohazard excep-ons. P2a â Repor-ng completeness Govern 1.5 (feedback & monitoring) §9.3 Management review inputs Art. 11(1)(d) Tes-ng & valida-on results Mandatory repor-ng of nulls and nega-ve findings; explicit an--selec-ve- repor-ng commitment. P2b â COI & access disclosure Govern 1.1 (roles & responsibili-es) §5.3 Organisa-onal roles Art. 52 Informa-on to downstream providers Disclosure of privileged access to models, data, or compute unavailable to external replicators. P2c â Provenance integrity Manage 4.1 (risk treatment documenta-on) §7.5 Documented informa-on control Art. 11(1)(b) Design & development process Tamper-evident logs for core lifecycle events; public correc-on channel with stated errata policy. P3 â Evidence- propor-onal claims Measure 2.6 (pre- deployment tes-ng) §8.4 AI system impact assessment Art. 55 Evalua-on & tes-ng obliga-ons Preregistered analysis plans; claim strength explicitly matched to quan-fied support; robustness checks propor-onate to stakes. P4 â Independent verifica-on Measure 2.7; Manage 3.2 (third- party tes-ng) §9.2 Internal audit Art. 55(1) Adversarial tes-ng At least one independent adversarial verifica-on for major safety claims; full replica-on packages as default. P5 â Artefact stewardship Govern 1.7 (documenta-on standards) §7.5 Documented informa-on; §8.3 AI system lifecycle Art. 11 Technical documenta-on archival DOIs and versioned releases; environment specs; hash manifests; long-term ownership and maintenance statements. P6 â Ins-tu-onalised scep-cism Measure 2.7 (independent evalua-on) §10.2 Con-nual improvement Art. 55(1) Red- teaming Red-team review as default for safety claims; published cri-que reports; career credit and boun-es for flaw-finding. P7 â Uncertainty accoun-ng Measure 2.3, 2.5 (metrics & monitoring) §9.1 Monitoring, measurement, analysis Art. 55(3) Post- market monitoring Living uncertainty register; structured quan-fica-on of sta-s-cal and epistemic unknowns; calibra-on audits; residual-risk statements. P8 â Decision relevance Manage 4.2 (risk communica-on) §6.2 AI policy; §8.2 Risk criteria Art. 52â53 Informa-on for deployers & users Explicit claim boundaries (âdoes/does not implyâ); decision-facing guidance with reliance condi-ons; update triggers. 6.2 Motivational X-Risk Essays A body of grey-literature essays has arHculated why extreme epistemic cauHon in AI is needed. For instance, Alex Flintâs âAI Risk for Epistemic Minimalistsâ (2021) disHls the AI existenHal risk argument into four premises: (i) highly capable AI is likely this century, (i) rapid shins in power historically 19 entail existenHal danger, (i) profit-driven, uncoordinated AI deployment increases risk, and (iv) prudence demands precauHon. Flintâs essay compellingly argues that we must take AI risk seriously, but it leaves open how researchers should conduct themselves to jusHfy and audit safety claims. Our proposed Code can be seen as an answer to that lacuna: Principles P1âP8 turn Flintâs abstract call for precauHon into concrete research duHes â e.g. radical transparency (P1), adversarial red-teaming (P6), reflexive uncertainty audits (P7) â providing the procedural backbone that Flint declines to specify. In other words, given the premises that drive urgency, we offer a systemaHc method to ensure our knowledge claims about safety are solid and verifiable. 6.3 Philosophical grounding: from Kuhnâs values to ECAISA Kuhn holds that scienHfic communiHes choose among frameworks by appealing to shared but imprecise values including accuracy, (internal/external) consistency, scope, simplicity, and fruiulness, rather than to any single decision procedure (Kuhn 1970; 1977). ECAISA makes these values auditable under fronHer-AI condiHons. P3 (Evidence-ProporHonal Claims) and P7 (Reflexive Uncertainty AccounHng) render accuracy and consistency operaHonal via quanHfied support, calibraHon, and explicit uncertainty registers; P4 (Reproducibility & Adversarial VerificaHon) and P6 (Organised ScepHcism) insHtuHonalise severe tests that extend demonstrated scope while guarding against spurious âfruiulnessâ; P1 (Radical Transparency) and P5 (FAIR Stewardship) enable communal uptake by making problem-solving pracHces legible; and P8 (Epistemic Hygiene) disciplines simplicity by separaHng warranted parsimony from Hdy narraHve. Because alignment adds an asymmetric objecHve, i.e., minimising tail risk, we extend Kuhnâs value set with a domain-specific criterion. This includes safety margin under worst-case uncertainty, made acHonable through adversarial evaluaHon and safety-case argumentaHon. Thus, ECAISA does not replace Kuhnâs values, instead it renders them verifiable for a field where incommensurability and high stakes otherwise reward premature consensus. 6.4 Philosophy of Science for Alignment In the AI alignment forum sphere, Nora Ammannâs ongoing LessWrong sequence âThoughts in the Philosophy of Science of AI Alignmentâ (Ammann 2023) has been probing foundaHonal epistemic quesHons: What counts as evidence in alignment research? How should progress be measured? Which scienHfic metaphors are apt or misleading? Ammannâs posts clarify concepts (e.g. comparing alignment research to search vs. verificaHon paradigms) and offer heurisHcs, but as Ammann notes, the sequence âintenHonally stops short of formal norm-seOng.â Our framework directly builds on that groundwork by codifying Ammannâs insights into enforceable norms. Ammann also discusses the importance of preregistraHon and avoiding unfalsifiable claims â we incorporate that as P1 (preregistraHon and openness) and P3 (evidence proporHonality). She advocates âepistemic humilityâ and organized scepHcism; we formalize those under P6 and P7. Furthermore, we align these norms with governance âhooksâ like NISTâs AI Risk Management Framework 1.0 and the emerging ISO 42001 standard (which Ammann did not address), thus moving from reflecHve philosophy to acHonable policy. In short, whereas Ammann provides conceptual raHonale, ECAISA provides the rules and processes to implement those principles in daily research. 6.5 Sociology of the AI Safety Field On the empirical side, Ahmed et al. (2024) conducted an ethnography Htled âFieldbuilding and the Epistemic Culture of AI Safety,â analysing how the AI safety community (parHcularly rooted in effecHve altruism and longtermism) has formed an âepistemic communityâ. They observed forums, forecasHng plaorms, and prize compeHHons, idenHfying mechanisms that give the community influence (e.g. close-knit networks, shared jargon) but also potenHal epistemic monoculture risks. While this study richly describes how the community self-organizes and the social reinforcement at play, it does not prescribe minimum research-quality thresholds or norms. 20 Our Code supplies those missing safeguards: for example, Ahmed et al. note the communityâs heavy reliance on trust and reputaHon; our principles inject transparent verificaHon requirements to counteract any insularity (P1, P4). They menHon the outsized policy sway of a relaHvely small group; our Code, by requiring explicit documentaHon and red-team disclosures, provides a way to earn and check that influence. We operaHonalise the latent values Ahmed et al. idenHfy turning informal habits (like sharing ideas on forums) into formal norms (like mandatory public sharing of research arHfacts). By doing so, we aim to miHgate the monoculture risk: even if a Hght community forms, its claims are legible and challengeable by outsiders, thus reducing closed-group dynamics. 6.6 High-Reliability and Safety Engineering Parallels There is extensive literature in safety engineering and high-reliability organizaHons (HROs) that indirectly speaks to our aims. For example, Weick and Sutcliffeâs work on HROs (Weick et. al 2015) outlines cultural principles like preoccupaHon with failure, reluctance to simplify, and commitment to resilience â hallmarks of organizaHons (like nuclear power plants or air traffic control centres) that consistently avoid catastrophe. Our proposed norms resonate strongly with these. A preoccupaHon with failure is exactly what our âfailure vs. capabilityâ asymmetry and Principle P6 (insHtuHonalized scepHcism) encourage â never resHng on laurels, always asking âWhat might we have missed?â (Pozzobon et al. 2023). P7 resists oversimplificaHon: it requires engaging with complexity and uncertainty instead of assuming unknowns away. Commitment to resilience and deference to experHse in crises maps to ensuring that when surprises occur, we rapidly adapt and involve whoever has the relevant knowledge (analogous to bringing in outside experts for criHque, per P8). While HRO theory is about organizaHonal structure and culture, its ethos supports our argument that extreme rigor and mindfulness can prevent disasters. Our contribuHon is to translate those general HRO principles into research epistemology: for example, requiring specific documentaHon, adversarial tesHng, and mulH-layered safety checks in alignment research (as our Code does), as opposed to broad organizaHonal advice. AddiHonally, risk governance scholars have argued that when stakes are high, facts uncertain, and decisions urgent, science must adopt new norms like extended peer review and the precauHonary principle. This is the essence of âpost-normal scienceâ as described by Funtowicz and Ravetz (1993). Our Code can be seen as implemenHng a post-normal science approach for AI: we invite âextended peersâ (via public audits, community criHque, and red-team exercises) and essenHally enshrine precauHon (Shaw 2009). P3 and P6 ensure claims are extraordinarily well-supported before being accepted as resolved. This connecHon situates our work in a broader scholarly trend of adapHng scienHfic pracHce for unprecedented, high-stakes challenges (climate science is onen cited as another domain requiring such epistemic shins). In summary, we build a bridge between classic safety engineering rigor and these newer paradigms by concretely specifying how to achieve evidence of absence in AI research (a theme also developed in recent commentary at University of No;ngham Blogs 1 ). 6.7 Methodological Tools for Uncertainty and Assurance On the technical front, there are surveys like Wang et al. (2020) âFrom Aleatoric to Epistemic: A Review of Uncertainty QuanHficaHon (UQ) Methods in Deep Learning,â which catalogue various techniques such as probabilisHc modelling, Bayesian neural nets, ensembles, etc. relevant to AI safety. They provide the toolbox for measuring uncertainty but, as the authors note, their review is methodological, not normaHve: it tells how to measure uncertainty, not when or whether researchers should report it. Our Codeâs Principle P7 (Reflexive Uncertainty AccounHng) essenHally takes those UQ techniques and makes their use compulsory in alignment research. We integrate UQ into a broader duty: not only must you use such tools, but you must also iterate and disclose 1 blogs.nobngham.ac.uk 21 uncertainty results (e.g. perform calibraHon over Hme, use disclosure ladders to communicate uncertainty publicly). We push UQ from an opHonal enhancement to a âstanding epistemic dutyâ. When empirical signal is thin, teams should deploy structured expert judgment protocols, e.g., Cookeâs performance-weighted âClassical Modelâ, to elicit, score, and aggregate expert probabiliHes, with seed quesHons and calibraHon diagnosHcs reported alongside results (Cooke 1991). Similarly, the literature on formal verificaHon and assurance cases (common in sonware safety) informs our approach. For instance, safety-criHcal industries onen use safety cases (structured arguments backed by evidence that a system is safe). One could view our enHre ECAISA code as guidelines to produce a robust âepistemic safety caseâ for an AI system including evidence of tesHng, adversarial challenges, uncertainty bounds, etc. Our work complements recent proposals in AI policy that suggest requiring explicit assurance documentaHon for advanced AI (as hinted in the UKâs AI White Paper and other governance discussions). We provide concrete content that would go into such documentaHon if our norms were adopted. 6.8 Epistemic Culture Cri6ques and Alterna6ves There have also been meta-level criHques of how the AI alignment community handles knowledge and dissent. Williams (2025a), warns that stylisHc conformity and deference to status hierarchies in the community could âthrople paradigm-breaking ideas.â He calls for cross-disciplinary stress tests and even proposes âepistemic overfiOng audits.â In follow-up work, Williams (Williams 2025b) examines epistemic closure dynamics, how certain ideas can become invisible due to community blind spots, and advocates for decentralized, diverse approaches (e.g. involving a much wider collecHve intelligence in alignment research). Our Code tackles similar hazards but from within exisHng insHtuHons: by Hghtening transparency (P1), integrity (P2), and reproducibility (P4), we aim to lower barriers to entry and make it easier for outsiders or independent voices to scruHnize and contribute to alignment research, thereby miHgaHng insularity. Williamsâs proposals (redesigning the enHre knowledge-producHon system) and ours (strengthening current research processes) are complementary: his approach is more radical and long-term, whereas ours can be immediately deployed by journals, conferences, and funding agencies without waiHng for a new insHtuHonal paradigm. AddiHonally, we note efforts in the open science and metascience communiHes. For instance, the NaHonal Academiesâ report Fostering Integrity in Research (NaHonal Academies of Sciences, Engineering, and Medicine 2017) which emphasise many of the same core values: objecHvity, honesty, openness, accountability, fairness, and stewardship. Our work aligns with these ideals but goes a step further in tailoring them to an AI context where adversaries and extreme tail risks exist factors that general research integrity frameworks do not explicitly consider. Lastly, at the policy level, the Asilomar AI Principles (Future of Life InsHtute, 2017) included a broad guideline on âResearch Cultureâ staHng that a culture of cooperaHon, trust, and transparency should be fostered among AI. Our proposal operaHonalizes that broad principle with specific measures. In a sense, ECAISA can be seen as detailing what it concretely means to have a transparent and trustworthy research culture for high-stakes AI. We move from aspiraHonal statements (like Asilomarâs call for a posiHve research culture) to defined standards against which compliance can be evaluated. 7. Discussion and Future Direc0ons ImplemenHng ECAISA is not cost-free. The tangible costs are documentaHon, packaging, independent verificaHon, and the coordinaHon Hme needed to run an infohazard review or a red-team exchange. We regard this overhead as the price of greater auditability and earlier error detecHon and note that high-reliability domains rouHnely treat comparable overhead as a cost of doing the work rather than a surcharge on it. Two pracHcal tensions nevertheless warrant discussion: the tension between 22 transparency and informaHon-hazard or commercial-confidenHality constraints, addressed by the disclosure ladder (Table 5) and the infohazard adjudicaHon protocol (§5.3); and the tension between adopHon cost and adopHon benefit under compeHHve condiHons, addressed in the adopHon- dynamics discussion below. Retrospec.ve rubric audit. To test whether the rubric in Table 4 is applicable beyond principle, we conducted a retrospecHve rubric audit on eight influenHal alignment papers drawn across four sub- areas (preference learning / RLHF; red-teaming and safety evaluaHon; mechanisHc interpretability; robustness and adversarial ML). Scoring was performed against the codebook in Appendix A, using only the paper text, appendices, and author-linked artefacts as evidence. The audit is explicitly an instrument-feasibility demonstraHon, not a field-wide audit: the sample is non-random, non- representaHve, and coded primarily by one author. A second coder independently scored a subset of three papers to give 24 doubly coded cells, yielding percent agreement 0.83 and weighted Cohenâs Îș = 0.79 (quadraHc weights, ordinal 0/1/2); these figures are consistent with substanHal agreement but have a necessarily wide confidence interval given the small overlap, and should be read as evidence that the rubric is scorable beyond a single author rather than as mature psychometric validaHon of the instrument. The eight audited papers were drawn across four sub-areas: preference learning and RLHF (ChrisHano et al. 2017; SHennon et al. 2020; Ouyang et al. 2022 [InstructGPT]; Bai et al. 2022 [ConsHtuHonal AI]), red-teaming and safety evaluaHon (Perez et al. 2022; OpenAI 2023 [GPT-4 Technical Report]), mechanisHc interpretability (Elhage et al. 2021), and robustness and adversarial ML (Madry et al. 2018). The selecHon was breadth-oriented rather than complete: it captures research subcultures that differ in their epistemic convenHons but is not a representaHve sample of any one of them. Inter-rater reliability is reported in Table 8; group-level means by sub-area are reported in Table 9. We disclose that the audit was conducted in-house by the author team. We treat this as an in-group reliability check, not as an external audit, and Stage 2 of the validaHon roadmap below specifically calls for an externally coded sample to test reliability beyond the author group. By the same standard ECAISA applies to the work it audits (P2b, P8), this access asymmetry is itself a feature of the evidence base and is reported here so that readers can weigh the Îș and group-mean figures accordingly. Table 8. Inter-rater reliability for the retrospecAve rubric audit (pilot overlap). Metric Value Papers double-coded (n) 3 Coded cells (papers Ă principles) 24 Percent agreement (all cells) 0.83 Weighted Cohenâs Îș (quadra:c weights, ordinal 0/1/2) 0.79 Table 9. RetrospecAve rubric audit: mean score per principle within sub-area groups (scale: 0 absent, 1 parAal, 2 meets). Group n P1 P2 P3 P4 P5 P6 P7 P8 RLHF / preference learning 4 0.75 0.00 1.00 0.75 0.50 0.00 0.00 1.00 Red-teaming / safety evalua-on 2 0.00 0.00 1.00 0.00 0.00 1.50 0.00 1.00 Mechanis-c interpretability 1 1.00 0.00 1.00 1.00 1.00 0.00 0.00 1.00 Robustness / adversarial ML 1 1.00 0.00 1.00 1.00 0.00 1.00 0.00 1.00 Three narrow conclusions from the audit are worth staHng. First, the rubric discriminates among pracHces rather than assigning uniformly high or uniformly low scores across all principles: evidence- proporHonal claims (P3) score consistently higher than full reporHng (P2), independent verificaHon 23 (P4), or uncertainty accounHng (P7), which suggests the criteria are sensiHve to real differences in pracHce rather than uniformly strict or lenient. Second, paperns are sub-community-specific in substanHvely intelligible ways: red-teaming papers score higher on P6 by design, while artefact- release pracHces under P1 and P5 vary across sub-communiHes. Third, the strongest ECAISA obligaHons, specifically independent adversarial verificaHon (P4) and structured uncertainty accounHng (P7), score near zero across all groups audited, idenHfying these as priority targets for adopHon incenHves. Three threats to validity warrant acknowledgement: sampling bias (papers were selected illustraHvely), coder subjecHvity (small reliability overlap), and construct validity (the rubric measures observable epistemic pracHces, which may track but do not equal research quality or safety impact). Whether higher rubric scores predict beper downstream outcomes is an empirical quesHon for later validaHon stages. Valida.on roadmap. The retrospecHve audit establishes only instrument feasibility. A stronger empirical case for ECAISA requires a staged validaHon programme. Stage 1 (fixed-sample reliability extension, near-term): expand double-coding within the exisHng eight-paper sample, adding two or three further coders from different disciplinary backgrounds, and report bootstrapped weighted-Îș confidence intervals. Stage 2 (expanded inter-rater reliability, near-term): conduct a larger preregistered reliability study on a straHfied random sample of at least 30 alignment papers with two independent coders, publishing the annotaHon sheet and adjudicaHon log. Stage 3 (comparaHve case study, medium-term): select matched papers with and without ECAISA-relevant pracHces and assess whether higher-scoring papers differ on downstream indicators such as replicaHon success, post- publicaHon error correcHon, and whether the evidence base was successfully challenged or extended by subsequent work. Stage 4 (prospecHve adopHon trial, longer-term): work with a venue, funder, or laboratory to adopt ECAISA requirements for a defined cohort and compare outcomes against a control cohort under exisHng norms, pre-registering the protocol, sampling frame, and analysis plan. Stages 3 and 4 provide the causal evidence needed to jusHfy insHtuHonal adopHon at scale; unHl they are run, ECAISA should be treated as a well-specified proposal with instrument feasibility demonstrated, not a validated intervenHon. 7.1 An6-gaming mechanisms A recurring concern about any checklist-based regime is that targets become gamed once they carry incenHves. The normalisaHon-of-deviance literature (Vaughan 1996) shows that well-intenHoned compliance regimes can calcify into box-Hcking rituals in which the form of the pracHce is preserved while its substance erodes. ECAISA does not claim to eliminate this risk. What it does is raise the expected cost of hollow compliance through a bundle of seven mutually reinforcing mechanisms, each designed to target a specific failure mode of checklist ritualism. (1) Randomised artefact spot-audits. Venues and funders sample accepted papers and funded projects for post-acceptance checks against stated artefacts, and publish aggregate outcomes. SelecHon is random, not Hed to controversy, so the expected cost of fabricaHng an artefact scales with the probability of selecHon. (2) LoMery-based independent verifica.on. For T2 and T3 claims, a pre-commiped fracHon must undergo independent reproducHon or adversarial verificaHon by a third party, with the specific claims selected by lopery rather than by venue discreHon. Lopery-based selecHon prevents a lab from avoiding verificaHon by posiHoning its strongest claims as lower-Her. (3) Release signing and hash manifests. Authors publish a version tag and a machine-readable manifest of file hashes for key artefacts. Subsequent undisclosed modificaHon of an artefact is therefore detectable, which ensures that the version subjected to spot-audit is the version originally released. 24 (4) Public correc.on channel. A public issue tracker and a stated errata policy are maintained for the replicaHon package. ParHes can file material errors as public issues; the authorâs response and update history become part of the audit record. This creates an ongoing reputaHonal surface that makes it costly to ignore known defects. (5) Logged disclosure excep.ons. For any restricHon applied under the disclosure ladder (L1âL3), the commipee records what is withheld, why, who can access it, and a re-evaluaHon date or condiHon. These logs are public even when the substanHve restricted content is not. This prevents âinfohazardâ from becoming an unapributable discreHon. (6) Red-team report minimum fields. Red-team reports are required to include a structured set of minimum fields: threat model, apack constraints, success criteria, and a public summary, even when full apack traces are controlled-access. A red-team report that omits these fields is itself a red flag under the scoring rubric. (7) Access-asymmetry disclosure. Authors disclose evaluaHon privileges that are not available to external replicators: privileged model access, proprietary compute, non-public evaluaHon harnesses. This lets reviewers and downstream users weight the evidence appropriately rather than treaHng all claims as equally replicable. Goodhartâs law and residual gaming risk. We do not claim these mechanisms eliminate gaming. They are designed to beat two specific baselines. Against no checklist at all, ECAISA makes epistemic pracHces observable, which is a precondiHon for any form of accountability. Against checklist-only compliance â the regime most vulnerable to Goodhart effects â the inclusion of adversarial verificaHon (mechanisms 1â2) and structured uncertainty registers (P7) shins the compliance target from easily fabricated formal outputs to substanHve epistemic pracHces that are harder to fake convincingly. An uncertainty register that lists no failure modes, for instance, is itself a red flag under P7âs scoring criteria and would be flagged by a competent spot-auditor. The residual risk is that well-resourced actors may invest in producing convincing but hollow compliance artefacts. We acknowledge this possibility and note two miHgaHons. First, the lopery- based independent verificaHon (mechanism 2) means that the expected cost of fabricaHon scales with the probability of selecHon, which venues can calibrate upward for T3 claims. Second, the public correcHon channel (mechanism 4) creates an ongoing reputaHonal surface: any researcher can file a public issue on a claimed artefact, creaHng a cost for authors whose compliance turns out to be hollow. These mechanisms do not guarantee substanHve compliance, but they make non-compliance observable and costly, which is the realisHc target for any governance instrument operaHng in a compeHHve research environment. 7.2 Adop6on dynamics and the coordina6on problem Voluntary adopHon of ECAISA by individual laboratories runs into a structural problem. Unilateral adopHon imposes costs â disclosure overhead, adversarial verificaHon, documentaHon, the coordinaHon Hme to run infohazard reviews â without reciprocal benefits when compeHtors do not adopt. Under the current deployment economics of fronHer AI, where speed-to-capability confers strategic advantage in funding, hiring, and market posiHon, a laboratory that unilaterally accepts hosHle replicaHon overhead for safety-relevant claims faces a first-mover disadvantage relaHve to compeHtors who do not. This is a classic coordinaHon failure, and any implementaHon outlook that ignores it underesHmates the gap between normaHve appeal and actual adopHon. Three realisHc adopHon paths reduce or bypass this unilateral disadvantage by shining the payoffs insHtuHonally rather than asking individual labs to accept the cost alone. (a) Venue-level mandates. Venues that host safety-relevant tracks can introduce artefact-based review criteria Hed to the ECAISA rubric for T2 and T3 claims: full replicaHon packages, adversarial 25 verificaHon reports, and uncertainty registers become prerequisites for acceptance rather than opHonal supplements. This flips the payoff structure: non-compliance becomes a rejecHon cost rather than a compeHHve advantage, and the marginal cost of compliance is paid by everyone who wants the venueâs imprimatur. Analogues include the CONSORT guidelines for clinical-trial reporHng and the preregistraHon requirements at registered-report venues; neither imposes compliance as a duty, but both make it strictly worthwhile by tying it to publicaHon access. (b) Funder mandates. Funding bodies that support safety-relevant research can earmark a fracHon of their safety-research budgets (we suggest in the range of 5â10%) for adversarial verificaHon grants, awarded compeHHvely to teams proposing credible refutaHon protocols for high-profile claims. This socialises the cost of independent verificaHon across the funding pool rather than asking each lab to find resources for hosHle replicaHon of its compeHtorsâ work. Similar arrangements have made external replicaHon programmes possible in experimental psychology and experimental economics. (c) Regulatory anchoring. Regulators and standards bodies can embed ECAISA-style hooks in the procedural layer of exisHng regulaHon: explicit research-phase evidenHary requirements in EU AI Act codes of pracHce, ISO/IEC technical reports on AI research quality management, and procurement frameworks that make compliance legible to downstream buyers. Once present at the procurement layer, compliance becomes a compeHHve asset rather than a burden, and labs that adopt ECAISA unilaterally recover the cost through the procurement channel. None of these paths eliminates first-mover disadvantage for a laboratory compeHng purely on capability. What they do is reduce the space in which pure capability compeHHon remains raHonal for safety-relevant claims specifically. The realisHc adopHon trajectory is therefore neither unilateral voluntary uptake (which the coordinaHon problem undermines) nor top-down mandate (which regulatory processes are too slow to produce in Hme), but a staged insHtuHonalisaHon through the three paths above, beginning with venue-level artefact requirements for T2 and T3 claims and funder support for independent adversarial verificaHon. Equity as infrastructure. A complementary consideraHon is equity. ECAISA increases documentaHon and verificaHon overhead; without insHtuHonal support, these demands risk entrenching incumbents with privileged access to compute, evaluators, and proprietary systems while excluding under- resourced researchers. AdopHon should therefore be paired with concrete support: earmarked replicaHon funding and bounHes for independent reproducHon and adversarial verificaHon of T2 and T3 claims; shared, community-governed evaluaHon harnesses with standardised reporHng; venue- provided artefact shepherding to help authors reach the P5 minimum without bespoke engineering capacity; and public-interest compute credits for reproducHons and red-teaming, prioriHsed for external validators. These measures treat verificaHon capacity as governance infrastructure rather than as an unfunded mandate on individual authors, and they help to break the coordinaHon failure from the bopom up as well as the top down. Moving forward, we outline several steps to promote adopHon of these epistemic norms and to evaluate their effecHveness: âą Pilot Projects and Case Studies: We encourage the community to trial the full P1âP8 stack on a few high-profile alignment research projects. For example, a large team working on a decepHve alignment benchmark or advanced AI sandbox experiment could formally adopt ECAISA for the projectâs duraHon and report on outcomes. Publishing detailed meta-analyses of these trials including benefits (e.g. errors caught, improvements in collaboraHon) and costs (extra Hme, resource needs) will build an empirical case for or against the codeâs recommendaHons. DemonstraHng real-world feasibility is key to broader buy-in. These pilots can also serve as exemplars or templates that others can emulate. 26 âą Incen.ves via Cer.fica.on and Publica.on Standards: We propose working with AI conferences, journals, and funding bodies to integrate ECAISA principles into their requirements. For instance, much like psychology journals now have âTransparency and Openness PromoHonâ (TOP) badges, alignment venues could establish an âAlignment Epistemic Rigorâ badge or checklist. A paper that adheres to all or most of P1âP8 (preregistered, all data open, adversarial review included, etc.) might be recognized formally, giving authors incenHve to do the extra work. Funding agencies could similarly prioriHze grants that commit to these pracHces (some grant programs already ask for reproducibility plans, this would be a logical extension). AddiHonally, an independent body could cerHfy labs or projects that follow the code (analogous to an ISO cerHficaHon but for research process), which would signal credibility to outsiders. Over Hme, if such cerHficaHons or norms become presHgious, they create a race to the top on rigor, countering the current race-to-publish incenHves. âą Integra.on into Governance and Policy: We will also advocate for these research-phase norms to be reflected in AI governance frameworks. Regulatory and standards bodies could incorporate ECAISA-like language into AI risk management guidelines. For example, research- phase audits or verificaHon standards could be appended to the EU AI Actâs codes of pracHce, ensuring that regulators explicitly value epistemic rigor alongside technical safety. ISO/IEC could produce a technical report on âAI research quality managementâ referencing these principles. By geOng alignment epistemics on the policy radar, we not only legiHmize the effort but might eventually require through law or funding mandates that safety-criHcal AI research follows stricter protocols. This would amplify the impact beyond what voluntary adopHon can achieve. UlHmately, formalizing epistemic norms is itself an experiment and an iteraHve process. We must apply the same scruHny to these principles that we demand for alignment claims. It is possible there are unintended consequences: for instance, could requiring preregistraHon and extreme proof slow down innovaHon or deter exploratory research? Might an overemphasis on process create bureaucraHc hurdles without proporHonate benefit? These are valid concerns that need empirical study. We envision feedback loops to refine the code: monitoring the fieldâs producHvity, creaHvity, and risk posture as norms change, and gathering community feedback regularly. For example, a meta- research study in a couple of years could compare alignment papers published under ECAISA guidelines vs. those that werenât, to see if differences in replicability or error rates emerge. If we find any of the principles are too onerous or not effecHve, we will adjust them. In this sense, ECAISA is a starHng framework grounded in Hmeless scienHfic virtues (honesty, transparency, scepHcism) but also adapHng to the urgent specifics of alignment (adversarial uncertainty, extreme tail risk). We also suggest exploring addiHonal dimensions the community raised that lie outside the scope of this paper. One is epistemic security protecHng the integrity of the research process from intenHonal manipulaHon (e.g. by malicious actors spreading misinformaHon or by AI systems themselves influencing our beliefs). Our principles indirectly help (because openness and verificaHon make it harder to insert false claims), but direct measures (like secure model evaluaHon environments or adversarial informaHon hygiene) might be needed. Another area is cross-disciplinary ferHlizaHon: collaboraHng with fields like cogniHve psychology (to understand and miHgate researcher biases) or economics of science (to design incenHves) could enrich the epistemic toolkit for alignment. Finally, while we have focused on research-phase norms, bridging the gap to deployment is important; how do these norms inform the transiHon from lab findings to real-world AI system governance? Future work could develop guidelines for how alignment research evidence should be translated into safety assurances for deployment (perhaps akin to medical translaHon from clinical trials to pracHce guidelines). 27 We argue that the adopHon of an Epistemic Code not as a one-off fix but as the beginning of a more reflecHve, self-regulaHng phase in the evoluHon of AI safety science. The communityâs willingness to experiment on itself (i.e. to criHcally evaluate and improve its own epistemic pracHces) will be a healthy indicator that we take the noHon of extreme rigor seriously. 8. Conclusion AI alignment is not merely a technical puzzle; it is fundamentally an epistemic challenge: how to extend reliable knowledge and guarantees into unprecedented, high-stakes domains. The stakes, potenHally the future of human civilizaHon demand research methods that are up to this task. We have argued that the alignment community should adopt a formal Epistemic Code encompassing transparency, integrity, evidence proporHonality, reproducibility, stewardship of research arHfacts, organized scepHcism, reflexive uncertainty accounHng, and epistemic hygiene. By adhering to these principles, researchers will document everything, challenge everything, and collaborate openly. In effect, we create a culture where any claim about safety is built on a mountain of openly scruHnized evidence and where doubt is systemaHcally welcomed to test claims. If the opHmisHc scenario materialises i.e. advanced AI integrates seamlessly without inflicHng harm, the quiet hero will be the discipline of our epistemic safeguards. As public-health experts know, when a pandemic never erupts, it is the unglamorous vaccinaHon programmes and surveillance protocols that deserve the credit. In safety-criHcal domains, non-events are evidence that precauHons worked. The same epistemic rigour refined for AI alignment could become a template for other fields operaHng under radical uncertainty such as climate geo-engineering, syntheHc-biology containment, even planetary-scale cybersecurity. By showing how to prosecute science when the downside risk is existenHal and the evidence sparse, alignment research can pioneer a broader paradigm of âhigh-uncertainty science.â In the final analysis, pursuing alignment is about expanding human knowledge under uncertainty but doing so responsibly. By internalising rigorous epistemic norms, the field maximises its credibility and its chances of truly solving the problem it addresses. The Alignment Epistemic Code offered here is a community-driven invitaHon to raise our own standards, to ensure that as we race to build safe AI, we are not cuOng corners in understanding what âsafeâ truly means. We believe such self-imposed discipline is not a hindrance but rather the foundaHon for trustworthy and lasHng progress toward a safer AI future. Acknowledgements The author likes to thank the editor and the anonymous reviewers for their detailed, constructive, and generous feedback throughout the review process. Their comments materially improved the clarity, scope, and evidential grounding of the paper. Appendix A: Rubric audit codebook This appendix specifies the coding protocol used to apply the rubric summarised in Table 4. It is intended to make the rubric reproducible, to clarify the evidenHary threshold for each score, and to specify conservaHve decision rules for ambiguous cases. 28 A.1 Unit of analysis and allowed evidence sources Unit of analysis. One âpaperâ consists of the main manuscript together with any officially linked appendices and supplementary materials, and any author-linked artefacts explicitly referenced in the paper at the Hme of coding. Allowed evidence sources. Coders consider only: (i) the PDF text (including appendices), (i) links embedded in the paper (URLs, DOIs), and (i) content reachable via those links without guessing addiHonal resources. Unpublished materials are not considered, pracHces are not inferred from author reputaHon, and informal claims (âavailable on requestâ) are not treated as evidence unless an access pathway is concretely specified. ConservaHve principle. If evidence for a criterion is ambiguous, missing, or not verifiable from the allowed sources, the lower score is assigned and the ambiguity is recorded in the jusHficaHon field. A.2 Scoring scale and general decision rules Each principle P1âP8 is coded on an ordinal scale: 0 (absent) = no substanHve evidence of the pracHce; 1 (parHal) = some evidence or parHal implementaHon, but falling short of the operaHonal threshold for âmeetsâ; 2 (meets) = saHsfies the operaHonal threshold defined in A.3 below. Tie-break rule: when deciding between 1 and 2, prefer 1 unless the âmeetsâ criterion is explicitly evidenced; when deciding between 0 and 1, prefer 0 unless substanHve evidence of the pracHce is present beyond generic prose. A.3 Per-principle criteria P1 â Transparency and disclosure. Score 0: no code, data, or evaluaHon details beyond prose; no stable links; no artefact access plan. Score 1: some artefacts provided or methods described plausibly supporHng reimplementaHon, but disclosure is incomplete or links are unstable. Score 2: sufficient artefacts released to enable meaningful scruHny, or a concrete disclosure protocol applied under the ladder (Table 5): (i) explicit list of withheld items, (i) jusHficaHon, (i) miHgaHons and access pathway, and (iv) a re-evaluaHon condiHon or date. P2a â Selec.ve-repor.ng preven.on. Score 0: selecHve reporHng evident; negaHve results and nulls absent; no limitaHons discussion. Score 1: some limitaHons noted but no demonstrated commitment to comprehensive reporHng. Score 2: explicit anH-selecHon pracHces including (i) failures and nulls reported alongside successes, (i) sufficient detail for others to challenge results, and (i) pre-specified primary outcomes or a clear post-hoc raHonale for reported metrics. P2b â Conflict-of-interest and access disclosure. Score 0: no funding or COI disclosure; no acknowledgement of access asymmetries. Score 1: basic funding or COI disclosed but privileged access affecHng replicability not noted. Score 2: full disclosure: (i) all funding and relevant conflicts named, and (i) any access asymmetries an independent replicator would not share are explicitly stated. P2c â Provenance and error-correc.on integrity. Score 0: no mechanism for post-publicaHon correcHon; no versioning or lifecycle logging. Score 1: some artefact versioning exists but no explicit correcHon channel or errata policy. Score 2: operaHonal correcHon infrastructure: (i) public issue tracker or errata page with stated update policy; and (i) for T3 claims, tamper-evident provenance for core lifecycle events where feasible. P3 â Pre-specifica.on and quan.fied evalua.on. Score 0: largely qualitaHve claims without clear metrics or baselines. Score 1: quanHtaHve evaluaHon exists but under-specified (unclear baselines, limited robustness checks). Score 2: measurable success criteria with (i) clear metrics and baselines aligned to claims, (i) sufficient experimental detail, and (i) at least one robustness-oriented element proporHonate to the claim. 29 P4 â Independent and adversarial verifica.on. Score 0: no replicaHon or independent verificaHon. Score 1: strong internal verificaHon but no independent replicaHon or adversarial audit documented. Score 2: central claim subjected to an independent check (one of): independent replicaHon by a separate team; structured third-party red-teaming with reportable traces; or a release package designed for adversarial verificaHon plus explicit invitaHon and support for external replicaHon, with evidence of uptake or a clearly defined protocol. P5 â Artefact stewardship. Score 0: no artefacts or only ephemeral links. Score 1: artefacts exist but lack stewardship features (no version tags, no stable archival link or DOI, no environment spec). Score 2: artefacts stewarded as citable, durable research objects: (i) versioned release, (i) stable idenHfier, (i) environment specificaHon, (iv) licence and provenance notes, and (v) maintenance or ownership statement. P6 â Red-teaming and adversarial evalua.on. Score 0: no explicit adversarial or misuse-oriented tesHng. Score 1: some adversarial tesHng but without a clear threat model, systemaHc method, or reproducible traces. Score 2: structured adversarial evaluaHon: (i) threat model and apack surface; (i) systemaHc test design and constraints; (i) clear reporHng of successes and failures; and (iv) reproducible artefacts or documented excepHons. P7 â Uncertainty accoun.ng and calibra.on. Score 0: no meaningful uncertainty accounHng beyond generic limitaHons. Score 1: uncertainHes discussed in a non-structured way but lacks an explicit register or decision-relevant calibraHon. Score 2: explicit uncertainty artefacts including at least (i) a structured uncertainty register (assumpHons, known unknowns, failure modes), (i) a residual-risk statement and plausible falsifiers, and (i) calibraHon or sensiHvity analyses where applicable. P8 â Ins.tu.onal posi.oning and decision relevance. Score 0: no discussion of how claims should be used or misused; no COI disclosure; no governance context. Score 1: basic disclosure and some normaHve posiHoning but no operaHonal guidance for how the result should inform decisions. Score 2: claims explicitly situated in an insHtuHonal decision context: (i) clear claim boundaries (âwhat this does and does not implyâ), (i) explicit COI, funding, and relevant access-asymmetry disclosure, and (i) concrete decision-facing guidance. A.4 Inter-rater reliability For the pilot applicaHon of this codebook, a subset of papers is double-coded by an independent coder aner a brief calibraHon exercise on a paper outside the audit. We report percent agreement across coded cells and weighted Cohenâs Îș with quadraHc weights for the ordinal 0/1/2 scale (Cohen 1960, 1968). Disagreements are adjudicated using the decision rules above, with the default rule âprefer the lower score unless the criterion is explicitly evidenced.â Reliability indicators from a pilot overlap should be interpreted as evidence of scorability, not as a mature psychometric validaHon of the instrument; we recommend expanding double-coding in any larger subsequent study. A.5 Coding workflow and justification requirements Workflow: (1) read the abstract, introducHon, and contribuHons to idenHfy what the paper asserts; (2) idenHfy author-linked artefacts and record accessibility and coding date; (3) code P1âP8 using the criteria above, recording a short jusHficaHon per code; (4) record any ambiguous case and the reason for conservaHve scoring. The recommended annotaHon sheet fields are: paper ID, venue and year, claim type, scores P1âP8, jusHficaHons per principle, coding date, and artefact-access notes. Appendix B: Coverage-matrix justification This appendix substanHates every filled cell in the coverage matrix (Figure 1) by staHng the specific obligaHon each principle imposes that closes the relevant dimensionâs failure mode. The appendix is 30 intended to convert the matrix from an asserted mapping into an auditable claim: for each (principle, dimension) pair where a coverage mark appears, Table 10 names the concrete epistemic failure being addressed and the obligaHon through which the principle addresses it. Cells not listed are intenHonally empty in the matrix. Table 10. Coverage-matrix jusAficaAon. Each row lists one filled cell in Table 3 and names the failure mode it addresses and the obligaAon imposed. Principle Dimension Failure mode addressed ObligaAon imposed by the principle P1 E1 Op:misa:on target Worst-case evalua:ons hidden behind average-case repor:ng. Preregister hypotheses and release evalua:on pipelines so worst-case tests cannot be silently dropped. P1 E2 Transparency Delayed or par:al disclosure allows hidden failure modes to persist. Release code, data, and checkpoints at study outset in FAIR repositories (subject to the disclosure ladder). P1 E5 Enforcement culture Delayed disclosure weakens external review. Public preregistra:on creates an auditable commitment that reviewers can check against. P2 E2 Transparency Selec:ve repor:ng and undeclared interests distort the risk record. Require repor:ng of nulls, COI disclosure, access-asymmetry disclosure, and tamper-evident provenance. P2 E3 Verifica:on Missing provenance and undisclosed access make independent verifica:on impossible. Sub-obliga:on P2c: public correc:on channel with stated errata policy; P2b access-asymmetry disclosure. P2 E5 Enforcement culture Selec:ve repor:ng and hidden COIs normalise low- integrity prac:ce. Publicly log repor:ng commitments and COI; sanc:on departures via the correc:on channel. P3 E1 Op:misa:on target Claim strength untethered from the stringency of tail-risk evidence. Bind claim strength to quan:fied support; preregister analysis plans; require robustness checks propor:onate to stakes. P3 E2 Transparency Overclaiming obscures the real eviden:al basis. Transparent metrics and baselines; explicit absence-of-evidence caveats. P3 E3 Verifica:on Weak effect sizes inflated into strong claims without robustness. At least one robustness-oriented element propor:onate to claim strength. P3 E4 Uncertainty Overclaiming rooted in under- quan:fied uncertainty. Require explicit uncertainty range alongside any headline effect size. P4 E1 Op:misa:on target Worst-case claims unchallenged by adversaries. Mandate at least one independent adversarial verifica:on for major safety claims. P4 E3 Verifica:on Single-team results ossify without independent replica:on. Full replica:on package; structured invita:on for external replicators. P4 E5 Enforcement culture Lack of replica:on incen:ves reinforces single-team epistemics. Venues and funders are asked to reward adversarial verifica:on as first-class work (see Implementa:on Outlook). P5 E2 Transparency Ephemeral artefacts cannot be audited a_er publica:on. DOIs, versioned releases, rich metadata, and long-term maintenance statements. P5 E3 Verifica:on Non-reproducible environments block replica:on aYempts. Environment spec (lockfile or container), hash manifests for cri:cal files. P5 E5 Enforcement culture Lack of shared artefact standards makes external scru:ny costly. Community-governed stewardship with reproducible build pipelines for T3 claims. P6 E3 Verifica:on Adversarial review is informal Make red-team review the default for 31 and unincen:vised. safety claims; publish cri:que reports as first-class contribu:ons. P6 E5 Enforcement culture Weak ins:tu:onal scep:cism allows normalisa:on of deviance. Separa:on-of-concerns in internal safety teams; career credit and boun:es for flaw-finding. P7 E4 Uncertainty Shallow uncertainty repor:ng; epistemic unknowns go unrecorded. Maintain a living uncertainty register with assump:ons, failure modes, falsifiers, and residual-risk statements. P7 E1 Op:misa:on target Ta i l-risk claims made without quan:fying what is unknown. Require explicit residual-risk quan:fica:on aligned to the tail-risk target. P8 E5 Enforcement culture Research used as compliance proxy outside its intended scope. State claim boundaries; give decision- facing guidance with reliance condi:ons and update triggers. Bibliography Ahmed, Shazeda, Kasia JaĆșwiĆska, Akash Ahlawat, Alexi Winecoff, and Michael Wang. 2024. âField Building and the Epistemic Culture of AI Safety.â First Monday 29 (4). Altman, Douglas G., and J. Mar:n Bland. 1995. âAbsence of Evidence Is Not Evidence of Absence.â BMJ 311 (7003): 485. Amodei, Dario, Chris Olah, Jacob Steinhardt, Paul Chris:ano, John Schulman, and Dan ManĂ©. 2016. âConcrete Problems in AI Safety.â arXiv:1606.06565. Ammann, Nora. 2023. âThoughts in the Philosophy of Science of AI Alignment.â Sequence of essays on LessWrong (ongoing). Armstrong, D. M. 1973. Belief, Truth and Knowledge. Cambridge: Cambridge University Press. Bai, Yuntao, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, et al. 2022. "Cons:tu:onal AI: Harmlessness from AI Feedback." arXiv:2212.08073. BarreY, Stephen, Carmen CĂąrlana, Francesca Gomez, Joseph Sherlock, Emma Shercliff, and Louise Tipping. 2025. "Assessing Confidence in Fron:er AI Safety Cases." arXiv preprint. Bean, Andrew M., Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, et al. 2025. "Measuring What MaYers: Construct Validity in Large Language Model Benchmarks." In Advances in Neural Informa:on Processing Systems 38 (Datasets and Benchmarks Track). OpenReview ID: mdA5lVvNcU. Bengio, Yoshua. 2024. âTes:mony before the Canadian Senate Standing CommiYee on Social Affairs, Science and Technology, 18 March 2024.â OYawa: Parliament of Canada. ,,,. 2025. âPrepared Remarks for the AAAI Presiden:al Panel on AI Risk.â Presented at AAAI 2025, Aus:n, TX, February 2025. Bloomfield, Robin, and Bev LiYlewood. 2010. âMul: Legged Safety Cases for So_ware Based Systems.â So;ware Engineering Journal 25 (6): 435â444. Bloomfield, Robin, and John Rushby. 2024. "Assurance of AI Systems From a Dependability Perspec:ve." arXiv:2407.13948. Bostrom, Nick. 2011. Ar>ficial General Intelligence and the Problem of Control. Unpublished lecture notes, University of Oxford. ,,,. 2014. Superintelligence: Paths, Dangers, Strategies. Oxford: Oxford University Press. Buhl, Marie Davidsen, Gaurav SeY, Leonie Koessler, Jonas SchueY, and Markus Anderljung. 2024. "Safety Cases for Fron:er AI." arXiv:2410.21572. 32 Camerer, Colin F., Anna Dreber, Felix Holzmeister, Teck Hua Ho, JĂŒrgen Huber, Magnus Johannesson, Michael Kirchler, et al. 2018. âEvalua:ng the Replicability of Social Science Experiments in Nature and Science between 2010 and 2015.â Nature Human Behaviour 2: 637â44. hYps://doi.org/10.1038/s41562-018-0399-z. Carlsmith, Joseph. 2022. âPower Seeking AI: An X Risk Analysis.â Open Philanthropy, May 2022. Chris:ano, Paul F., Jan Leike, Tom Brown, Miljan Mar:c, Shane Legg, and Dario Amodei. 2017. "Deep Reinforcement Learning from Human Preferences." In Advances in Neural Informa:on Processing Systems 30, 4299â4307. Cirillo, Pasquale, and Nassim Nicholas Taleb. 2020. âTail Risk of Contagious Diseases.â Nature Physics 16: 606â 613. Cohen, Jacob. 1960. "A Coefficient of Agreement for Nominal Scales." Educa:onal and Psychological Measurement 20 (1): 37â46. hYps://doi.org/10.1177/001316446002000104. Cohen, Jacob. 1968. "Weighted Kappa: Nominal Scale Agreement with Provision for Scaled Disagreement or Par:al Credit." Psychological Bulle:n 70 (4): 213â20. hYps://doi.org/10.1037/h0026256. Cohen, Jacob. 1988. Sta>s>cal Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale, NJ: Lawrence Erlbaum. Cooke, Roger M. 1991. Experts in Uncertainty: Opinion and Subjec>ve Probability in Science. New York: Oxford University Press. DARPA. 2019. âSystema:zing Confidence in Open Research and Evidence (SCORE) Program Overview.â Arlington, VA: Defense Advanced Research Projects Agency. De Angelis, Catherine D., Jeffrey M. Drazen, Frank A. Frizelle, CharloYe Haug, John P. Hoey, Drummond Rennie, and Harold C. Sox. 2004. âClinical Trial Registra:on: A Statement from the Interna:onal CommiYee of Medical Journal Editors.â New England Journal of Medicine 351 (12): 1250â1251. Department for Digital, Culture, Media & Sport (UK). 2023. Establishing a Pro-Innova>on Approach to AI Regula>on (White Paper). London: DCMS. Elhage, Nelson, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, et al. 2021. "A Mathema:cal Framework for Transformer Circuits." Transformer Circuits Thread, December 22, 2021. hYps://transformer-circuits.pub/2021/framework/index.html. European Commission. 2025. Code of Prac>ce on General Purpose Ar>ficial Intelligence (GPAI). Brussels: Directorate-General for Communica:ons Networks, Content and Technology, 9 July 2025. hYps://digital- strategy.ec.europa.eu/en/library/code-prac:ce-general-purpose-ai. Federal Avia:on Administra:on. 2022. Advisory Circular 25.1309-1B: System Design and Analysis. Washington, DC: FAA (dra_ update). Federal Avia:on Administra:on, Department of Transporta:on. 2024. âSystem Safety Assessments.â Federal Register 89 (166) (August 27): 68706â68735. hYps://w.federalregister.gov/d/2024-18511. Feynman, Richard P. 1986. âAppendix F: Personal Observa:ons on the Reliability of the ShuYle.â In Report of the Presiden>al Commission on the Space Shu†Challenger Accident, Vol. 2. Washington, DC: Government Prin:ng Office. Flint, Alex. 2021. "AI Risk for Epistemic Minimalists." LessWrong, August 22, 2021. h4ps://w.lesswrong.com/posts/8fpzBHt7e6n7Qjoo9/ai-risk-for-epistemic-minimalists. Future of Life Ins:tute. 2017. âAsilomar AI Principles.â Published 11 August 2017. hYps://futureoflife.org. Funtowicz, Silvio O., and Jerome R. Ravetz. 1993. âScience for the Post Normal Age.â Futures 25 (7): 739â755. Ganguli, Deep, Jared Kaplan, Tom Henighan, and Sam McCandlish. 2022. âRed Team Analysis of GPT-J 6B.â Anthropic Technical Report, October 2022. Gebru, Timnit, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal DaumĂ© I, and Kate Crawford. 2021. âDatasheets for Datasets.â Communica>ons of the ACM 64 (12): 86â92. hYps://doi.org/10.1145/3458723. 33 Goemans, Anouk, Marie Davidsen Buhl, Jonas SchueY, Tomek Korbak, Jessica Wang, Benjamin Hilton, and Geoffrey Irving. 2024. "Safety Case Template for Fron:er AI: A Cyber Inability Argument." arXiv:2411.08088. Gnei:ng, Tilmann, and Adrian E. Ra_ery. 2007. "Strictly Proper Scoring Rules, Predic:on, and Es:ma:on." Journal of the American Sta>s>cal Associa>on 102 (477): 359â378. hYps://doi.org/10.1198/016214506000001437. Goodfellow, Ian J., Jonathon Shlens, and Chris:an Szegedy. 2015. âExplaining and Harnessing Adversarial Examples.â In Proceedings of the 3rd Interna>onal Conference on Learning Representa>ons (ICLR 2015). GreenblaY, Ryan, Carson Denison, Benjamin Wright, Fabien Roger, and Monte MacDiarmid. 2024. âAlignment Faking in Large Language Models.â arXiv, December 20, 2024. hYps://doi.org/10.48550/arXiv.2412.14093. Gundersen, Odd Erik, and SigbjĂžrn Kjensmo. 2018. âState of the Art: Reproducibility in Ar:ficial Intelligence.â In Proceedings of the 32nd AAAI Conference on Ar>ficial Intelligence, 1644â1651. Hanson, Robin. 1995. âCould Gambling Save Science? Encouraging an Honest Consensus.â Social Epistemology 9 (1): 3â33. hYps://doi.org/10.1080/02691729508578768. Henderson, Peter, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. 2018. âDeep Reinforcement Learning That MaYers.â In Proceedings of the AAAI Conference on Ar>ficial Intelligence, 3207â 3214. Hilton, Benjamin, Marie Davidsen Buhl, Tomek Korbak, and Geoffrey Irving. 2025. âSafety Cases: A Scalable Approach to Fron:er AI Safety.â arXiv preprint, March. arXiv:2503.04744. hYps://arxiv.org/abs/2503.04744. Hoe:ng, Jennifer A., David Madigan, Adrian E. Ra_ery, and Chris T. Volinsky. 1999. âBayesian Model Averaging: A Tutorial.â Sta>s>cal Science 14 (4): 382â417. Hubinger, Evan, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and ScoY Garrabrant. 2019. âRisks from Learned Op:miza:on in Advanced Machine Learning Systems.â MIRI Technical Report, August 2019. Interna:onal Atomic Energy Agency. 2016. Safety of Nuclear Power Plants: Design. Safety Standards Series No. SSR-2/1 (Rev. 1). Vienna: IAEA. Interna:onal Organiza:on for Standardiza:on. 2019. ISO 14971:2019 , Medical Devices,Applica>on of Risk Management to Medical Devices. Geneva: ISO. Interna:onal Organiza:on for Standardiza:on / Interna:onal Electrotechnical Commission. 2023. ISO/IEC 42001:2023 , Informa>on Technology , Ar>ficial Intelligence , Management System. Geneva: ISO. Ioannidis, John P. A. 2005. "Why Most Published Research Findings Are False." PLoS Medicine 2 (8): e124. hYps://doi.org/10.1371/journal.pmed.0020124. Katz, Guy, Clark BarreY, David Dill, Kyle Julian, and Mykel Kochenderfer. 2017. âReluplex: An Efficient SMT Solver for Verifying Deep Neural Networks.â In Computer Aided Verifica>on, 97â117. Berlin: Springer. Knub, Reto, Flavio L. Meehl, Claudia M. Goodess, and Tim Stocker. 2010. âChallenges in Combining Projec:ons from Mul:ple Climate Models.â Journal of Climate 23 (10): 2739â2758. Kuhn, Thomas S. 1970. The Structure of Scien>fic Revolu>ons. 2nd ed., enlarged. Chicago: University of Chicago Press. Kuhn, Thomas S. 1977. âObjec:vity, Value Judgment, and Theory Choice.â In The Essen>al Tension: Selected Studies in Scien>fic Tradi>on and Change, 320â39. Chicago: University of Chicago Press. Langley, Pat. 2019. âPromo:ng Openness and Credibility in AI Research.â AI Magazine 40 (4). Langosco, Daniele, Jack Koch, Laura R. Clark, Jonathan Uesato, Lawrence Chan, et al. 2022. âGoal Misgeneraliza:on in Deep Reinforcement Learning.â In Proceedings of the 39th Interna>onal Conference on Machine Learning (ICML 2022). Leveson, Nancy G. 2012. Engineering a Safer World: Systems Thinking Applied to Safety. Cambridge, MA: MIT Press. Lempert, Robert J., Steven W. Popper, and Steve C. Bankes. 2003. Shaping the Next One Hundred Years: New Methods for Quan>ta>ve, Long-Term Policy Analysis. Santa Monica, CA: RAND Corpora:on. 34 LiYlewood, Bev, and Lorenzo Strigini. 1993. âValida:on of Ultra High Dependability for So_ware Based Systems.â Communica>ons of the ACM 36 (11): 69â80. Madry, Aleksander, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. âTowards Deep Learning Models Resistant to Adversarial AYacks.â In Proceedings of the Interna>onal Conference on Learning Representa>ons (ICLR 2018). Mayo, Deborah G. 2018. Sta>s>cal Inference as Severe Tes>ng: How to Get Beyond the Sta>s>cs Wars. Cambridge: Cambridge University Press. McCaslin, Tegan, Jide Alaga, Samira Nedungadi, Seth Donoughe, Tom Reed, Rishi Bommasani, Chris Painter, and Luca RigheO. 2025. "STREAM (ChemBio): A Standard for Transparently ReporHng EvaluaHons in AI Model Reports." arXiv:2508.09853. hpps://doi.org/10.48550/arXiv.2508.09853. Merton, Robert K. 1942. âA Note on Science and Technology in a Democra:c Order.â Journal of Legal and Poli>cal Sociology 1: 115â126. ,,,. 1973. The Sociology of Science: Theore>cal and Empirical Inves>ga>ons. Chicago: University of Chicago Press. Mitchell, Margaret, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. âModel Cards for Model Repor:ng.â In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT â19), 220â29. New York: Associa:on for Compu:ng Machinery. hYps://doi.org/10.1145/3287560.3287596. MunafĂČ, Marcus R., Brian A. Nosek, Dorothy V. M. Bishop, Katherine S. BuYon, Christopher D. Chambers, Nathalie Percie du Sert, Uri Simonsohn, et al. 2017. âA Manifesto for Reproducible Science.â Nature Human Behaviour 1 (0021). hYps://doi.org/10.1038/s41562-016-0021. Narayanan, Arvind, and Sayash Kapoor. 2023. "Evalua:ng LLMs Is a Minefield." Knight First Amendment Ins:tute, Columbia University. hYps://knightcolumbia.org/blog/evalua:ng-llms-is-a-minefield. Na:onal Academies of Sciences, Engineering, and Medicine. 2017. Fostering Integrity in Research. Washington, DC: The Na:onal Academies Press. hYps://doi.org/10.17226/21896. â. 2019. Reproducibility and Replicability in Science. Washington, DC: The NaOonal Academies Press. h4ps://doi.org/10.17226/25303. Na:onal Diet of Japan. 2012. The Official Report of the Fukushima Nuclear Accident Independent Inves>ga>on Commission. To k y o : T h e C o m m i s s i o n . Na:onal Ins:tute of Standards and Technology. 2023. Ar>ficial Intelligence Risk Management Framework (AI RMF 1.0). Gaithersburg, MD: NIST. NRC (U.S. Nuclear Regulatory Commission). 2017. NUREG-1855, Rev. 1: Guidance on the Treatment of Uncertain>es Associated with PRAs in Risk-Informed Decision Making. Washington, DC: NRC. ,,,. 2018. Regulatory Guide 1.174, Rev. 3: An Approach for Using Probabilis>c Risk Assessment in Risk-Informed Decisions on Plant-Specific Changes to the Licensing Basis. Washington, DC: NRC. OpenAI. 2023. "GPT-4 Technical Report." arXiv:2303.08774. hYps://arxiv.org/abs/2303.08774. Ouyang, Long, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, et al. 2022. "Training Language Models to Follow Instruc:ons with Human Feedback." In Advances in Neural Informa:on Processing Systems 35, 27730â27744. Perrow, Charles. 1999. Normal Accidents: Living with High Risk Technologies. 2nd ed. Princeton, NJ: Princeton University Press. Perez, Ethan, Douwe Kiela, Kyunghyun Cho, et al. 2022. âRed Teaming Language Models with Language Models.â arXiv:2210.10814. Pineau, Joelle, Koustuv Sinha, GeneviĂšve Fried, et al. 2021. âImproving Reproducibility in Machine Learning Research: A Report from the NeurIPS 2019 Reproducibility Program.â Journal of Machine Learning Research 22 (238): 1â20. 35 Popper, Karl. 1959. The Logic of Scien>fic Discovery. London: Hutchinson. Pozzobon, L. D., J. Lam, E. Chimonides, B. Perkins Meingast, and W. S. Luk. 2023. âAdop:ng High Reliability Organiza:on Principles to Lead a Large Scale Clinical Transforma:on.â Healthcare Management Forum 36 (4): 241â45. hYps://doi.org/10.1177/08404704231162785. Presiden:al Commission on the Space ShuYle Challenger Accident. 1986. Report to the President. Washington, DC: Government Prin:ng Office. Raff, Edward. 2019. âA Step Toward Quan:fying Independently Reproducible Machine Learning Research.â In Advances in Neural Informa>on Processing Systems 32, 2195â2205. Raji, Inioluwa Deborah, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. 2020. "Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Audi:ng." In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (FAT* '20), 33â44. New York: ACM. hYps://doi.org/10.1145/3351095.3372873. Reuel, Anka, Amelia Hardy, Chandler Smith, Max Lamparth, Malcolm Hardy, and Mykel J. Kochenderfer. 2024. "BeYerBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Prac:ces." In Advances in Neural Informa:on Processing Systems 37 (Datasets and Benchmarks Track). arXiv:2411.12990. RTCA, Inc. 2011. DO-178C: So;ware Considera>ons in Airborne Systems and Equipment Cer>fica>on. Washington, DC: RTCA, Inc. Russakovsky, Olga, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, et al. 2015. âImageNet Large Scale Visual Recogni:on Challenge.â Interna>onal Journal of Computer Vision 115 (3): 211â252. Russell, Stuart, Daniel Dewey, and Max Tegmark. 2015. âResearch Priori:es for Robust and Beneficial Ar:ficial Intelligence: An Open LeYer.â AI Magazine 36 (4): 105â114. Shevlane, Toby, Sebas:an Farquhar, Ben Garfinkel, Mary Phuong, Jess WhiYlestone, and Jade Leung. 2023. âModel Evalua:on for Extreme Risks.â arXiv:2305.15324. Shaw, Christopher. 2009. âThe Dangerous Limits of Dangerous Limits: Climate Change and the Precau:onary Principle.â The Sociological Review 57 (2, suppl.): 103â123. hYps://doi.org/10.1111/j.1467-954X.2010.01888.x. S:ennon, Nisan, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Chris:ano. 2020. "Learning to Summarize from Human Feedback." In Advances in Neural Informa:on Processing Systems 33, 3008â3021. Te t l o c k , P h i l i p E . , a n d D a n G a r d n e r. 2 0 1 5 . Superforecas>ng: The Art and Science of Predic>on. New York: Crown. Turner, Barry A. 1976. â The Organiza:onal and Interorganiza:onal Development of Disasters.â Administra>ve Science Quarterly 21 (3): 378â397. Vaughan, Diane. 1996. The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA. Chicago: University of Chicago Press. Wan, Alexander, Kevin Klyman, Sayash Kapoor, Nestor Maslej, Shayne Longpre, BeYy Xiong, Percy Liang, and Rishi Bommasani. 2025. "The 2025 Founda:on Model Transparency Index." arXiv:2512.10169. Wang, Alex, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. âGLUE: A Mul:-Ta s k B e n c h m a r k a n d A n a l ys i s P l aÂo r m fo r N at u ra l L a n g u a g e U n d e rsta n d i n g .â I n Proceedings of the 2019 Interna>onal Conference on Learning Representa>ons (ICLR). Wang, Alex, Yada Pruksachatkun, Nikita Nangia, et al. 2020. âSuperGLUE: A S:ckier Benchmark for General-Purpose Language Understanding Systems.â In Advances in Neural Informa>on Processing Systems 33, 3266â3280. Weick, Karl E., and Kathleen M. Sutcliffe. 2015. Managing the Unexpected: Sustained Performance in a Complex World. 3rd ed. Hoboken, NJ: Wiley. Wieringa, Maranke. 2020. "What to Account for When Accoun:ng for Algorithms: A Systema:c Literature Review on Algorithmic Accountability." In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (FAT* '20), 1â18. New York: ACM. hYps://doi.org/10.1145/3351095.3372833. 36 Wilkinson, Mark D., Michel Dumon:er, IJsbrand Jan Aalbersberg, Gabrielle Appleton, Myles Axton, Arie Baak, Niklas Blomberg, et al. 2016. âThe FAIR Guiding Principles for Scien:fic Data Management and Stewardship.â Scien>fic Data 3 (160018). hYps://doi.org/10.1038/sdata.2016.18. Williams, A. E. 2025a. âEpistemic Closure and the Irreversibility of Misalignment: Modeling Systemic Barriers to Alignment Innova:on.â arXiv 2504.02058. hYps://doi.org/10.48550/arXiv.2504.02058. Williams, A. E. 2025b. âExpanding AI and AI Alignment Discourse: An Opportunity for Greater Epistemic Inclusion.â PhilArchive. hYps://philarchive.org/rec/WILEAA-29.