Paper deep dive
Compliant But Unsatisfactory: The Gap Between Auditing Standards and Practices for Probabilistic Genotyping Software
Angela Jin, Alexander Asemota, Dan E. Krane, Nathaniel D. Adams, Rediet Abebe
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/14/2026, 2:26:26 AM
Summary
This paper investigates the gap between the intended outcomes of audit standards and their actual implementation, using ASB 018 (a standard for probabilistic genotyping software in forensic DNA analysis) as a case study. The authors find that while forensic labs can achieve technical compliance with ASB 018, the audits often fail to meet the standard's envisioned goals, such as establishing operational boundaries or testing the broader sociotechnical system. The study attributes these failures to design flaws in the standard, including vague language and undefined terms, and provides recommendations for designing more effective audit standards.
Entities (5)
Relation Signals (3)
American Academy of Forensic Sciences Standards Board â developed â ASB 018
confidence 100% ¡ developed and published in 2020 by the American Academy of Forensic Sciences Standards Board
ASB 018 â governs â Probabilistic Genotyping Software
confidence 100% ¡ ASB 018, a standard for auditing probabilistic genotyping software
U.S. Criminal Legal System â uses â Probabilistic Genotyping Software
confidence 100% ¡ software that the U.S. criminal legal system increasingly uses to analyze DNA samples
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:AI governance efforts increasingly rely on audit standards: agreed-upon practices for conducting audits. However, poorly designed standards can hide and lend credibility to inadequate systems. We explore how an audit standard's design influences its effectiveness through a case study of ASB 018, a standard for auditing probabilistic genotyping software -- software that the U.S. criminal legal system increasingly uses to analyze DNA samples. Through qualitative analysis of ASB 018 and five audit reports, we identify numerous gaps between the standard's desired outcomes and the auditing practices it enables. For instance, ASB 018 envisions that compliant audits establish restrictions on software use based on observed failures. However, audits can comply without establishing such boundaries. We connect these gaps to the design of the standard's requirements such as vague language and undefined terms. We conclude with recommendations for designing audit standards and evaluating their effectiveness.
Tags
Links
- Source: https://arxiv.org/abs/2604.10875v1
- Canonical: https://arxiv.org/abs/2604.10875v1
Trouble viewing inline? Open PDF directly â
Full Text
129,583 characters extracted from source content.
Expand or collapse full text
by Compliant But Unsatisfactory: The Gap Between Auditing Standards and Practices for Probabilistic Genotyping Software Angela Jin University of California, BerkeleyBerkeleyCaliforniaUSA angelacjin@berkeley.edu , Alexander Asemota University of California, BerkeleyBerkeleyCaliforniaUSA alexander.asemota@berkeley.edu , Dan E. Krane Wright State UniversityDaytonOhioUSA dan.krane@wright.edu , Nathaniel Adams Forensic Bioinformatic Services, Inc.FairbornOhioUSA adams@bioforensics.com and Rediet Abebe ELLIS Institute; Max Planck Institute for Intelligent Systems; and TĂźbingen AI CenterTĂźbingenGermany aim-admin@tue.ellis.eu (2026) Abstract. AI governance efforts increasingly rely on audit standards: agreed-upon practices for conducting audits. However, poorly designed standards can hide and lend credibility to inadequate systems. We explore how an audit standardâs design influences its effectiveness through a case study of ASB 018, a standard for auditing probabilistic genotyping softwareâsoftware that the U.S. criminal legal system increasingly uses to analyze DNA samples. Through qualitative analysis of ASB 018 and five audit reports, we identify numerous gaps between the standardâs desired outcomes and the auditing practices it enables. For instance, ASB 018 envisions that compliant audits establish restrictions on software use based on observed failures. However, audits can comply without establishing such boundaries. We connect these gaps to the design of the standardâs requirements such as vague language and undefined terms. We conclude with recommendations for designing audit standards and evaluating their effectiveness. algorithm audits, accountability, standards, criminal legal system, forensic software, DNA profiling â journalyear: 2026â copyright: câ conference: Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems; April 13â17, 2026; Barcelona, Spainâ booktitle: Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI â26), April 13â17, 2026, Barcelona, Spainâ doi: 10.1145/3772318.3791552â isbn: 979-8-4007-2278-3/2026/04â ccs: Social and professional topics Government technology policyâ ccs: Software and its engineering Software verification and validationâ ccs: Applied computing Law 1. Introduction Figure 1. A simplified depiction of forensic lab use of probabilistic genotyping software to produce evidence in a criminal case. The goal of forensic DNA testing is to determine whether a person of interest (e.g., the defendant) could be the source of some or all of the DNA recovered from a crime scene. To this end, (A) a lab analyst first prepares two physical DNA samples: an evidence sample (e.g., a swabbing of a shoe collected as evidence) and a reference profile. (B) The analyst then runs the PGS with these DNA samples as inputs, along with other inputs chosen by the analyst. The software outputs a likelihood ratio (LR)âa statistic that describes similarities observed between the evidence sample and reference profile. (C) Before the LR is presented as evidence at trial, a court may conduct an admissibility hearing to determine whether the evidence is sufficiently reliable to be presented in front of a jury. (D) If the judge deems the evidence admissible, the LR will be presented as evidence at trial. . Flow-chart depicting stages of the forensic DNA testing process and the introduction of the probabilistic genotyping software (PGS) likelihood ratio output as evidence in a criminal case. The flowchart has four stages: (A) DNA profile preparation, (B) Use of PGS to produce likelihood ratio, (C) A potential admissibility hearing, and (D) Presentation of LR as evidence at trial. The figure depicts four inputs to the probabilistic genotyping software: the crime scene DNA sample (i.e., evidence sample), the person of interestâs DNA sample (i.e., reference profile), the forensic lab analystâs estimate of the number of contributors, and other analyst inputs. The software outputs a likelihood ratio (LR), which the figure then depicts as being presented in front of a judge. If the judge deems the LR admissible, the LR is shown as trial evidence in front of a jury. There is a box depicting the PGS. There is also a larger box encompassing both the PGS and lab analyst with a dotted line depicting the broader PGS system that takes physical DNA samples as input and outputs a LR. Over the past decade, researchers and practitioners have developed an ecosystem of toolsâincluding guidelines, frameworks, and software toolkitsâto promote fairness, transparency, and accountability throughout the AI lifecycle (Madaio et al., 2020; Holstein et al., 2019; Lee and Singh, 2021; Raji et al., 2020; Mitchell et al., 2019; Abebe et al., 2022). Amidst growing calls for algorithm audits, recent work highlights that such tools can be crucial for ensuring these evaluations result in meaningful accountability (Ojewale et al., 2025; Abebe et al., 2022; Lam et al., 2024; Raji et al., 2022; Jin and Salehi, 2024). However, tools for responsible AI may fail to achieve their intended outcomes when deployed in complex organizational settings and practices (Wong et al., 2023; Berman et al., 2024; Deng et al., 2022; Rakova et al., 2021; Madaio et al., 2020; Schor et al., 2024). In response, a burgeoning body of work investigates how to design audit tooling that effectively helps audits create accountability (Ojewale et al., 2025; Madaio et al., 2020; Raji et al., 2020; Holstein et al., 2019; Lam et al., 2024; Raji et al., 2022; Abebe et al., 2022). As researchers and policymakers increasingly call for audit standards (i.e., agreed-upon best practices) (Goodman and Trehu, 2022; Costanza-Chock et al., 2022; Lam et al., 2024; Raji et al., 2022; Groves et al., 2024; Hub, 2024), we ask: How can the design of an audit standard undermine its effectiveness? We specifically examine how a standardâs formal requirements may enable audits that comply with the letter of the standard while failing to meet the goals envisioned by the standardâs developers. We explore this question through the lens of ASB 018111We study the first edition of ASB 018, which was developed and published in 2020 by the American Academy of Forensic Sciences Standards Board (ASB) (ANSI/ASB Standard 018, 1st Ed., 2020)., a standard for auditing probabilistic genotyping software (PGS) used by U.S. forensic laboratories to analyze DNA samples in criminal cases. ASB 018 provides a critical case study not only because it mirrors audit processes in other domains like hiring (Raji et al., 2022; Lam et al., 2024; Groves et al., 2024; Gerchick et al., 2025), healthcare (Obermeyer et al., 2019; Wu et al., 2021), and social services (Chouldechova et al., 2018) but also because forensic labs already cite ASB 018 compliance to convince judges that PGS outputs are sufficiently reliable to be presented as evidence in criminal trials (United States v. Anderson, 673 F. Supp. 3d 671 (21-CR-204), (M.D. Pa. 2023)). We investigate ASB 018âs design and effectiveness through two research questions: (RQ1) What gaps exist between the practices ASB 018 envisions and the practices it enables?222We study forensic labsâ audit practices by examining their representation in publicly available audit reports; see Section 4., and (RQ2) How does the design of ASB 018 facilitate these gaps? In this work, we analyzed five publicly available PGS audit reports spanning various lab sizes, software versions, and years. We first assessed each audit report for compliance with ASB 018âs formal requirements. After determining that all five audits could be interpreted as compliant, we qualitatively coded ASB 018 and its accompanying factsheet to identify its desired outcomes across five audit stages (Ojewale et al., 2025): Audit Scope, Standards Identification, Performance Analysis, Post-Audit Judgment, and Audit Communication. We then qualitatively coded each audit report to identify practices that misalign with the standardâs goals (RQ1). Finally, we conducted a second round of coding on ASB 018, this time focusing on the language and organization of its requirements to identify aspects of its design that allow unsatisfactory audits to remain technically compliant (RQ2). Our findings reveal that the analyzed audits are ASB 018 compliant, but unsatisfactory: audit practices consistently fall short of ASB 018âs envisioned goals across all five audit stages, yet still satisfy the standardâs formal requirements. For instance, while ASB 018 envisions testing the broader sociotechnical system, the audits primarily measure narrow software performance. Similarly, although the standard intends for audits to establish boundaries on software use based on failures observed during performance analysis, the audits avoid framing errors as failures and establish no such boundaries. Furthermore, despite the standardâs emphasis on documentation that supports third-party evaluation, the reports omit crucial details needed to verify audit claims. ASB 018 enables these gaps through design choices such as undefined concepts, requirements language that treats audit components as separate considerations to be addressed in isolation, and vague language calling for labs to âaddressâ and âconsiderâ that allow compliance with minimal to no testing. To conclude, we discuss how our study demonstrates that clear articulation of a toolâs goalsâeven from an external perspectiveâis essential for evaluating and improving audit infrastructure. Towards designing audit tools that more effectively deliver accountability, we caution that designing for compatibility with usersâ current perspectives and practices can undermine the toolâs effectiveness. To help navigate this tension, we provide recommendations for those designing standards to (1) evaluate standards against fine-grained articulations of desired outcomes to identify where increased specificity is needed, using frameworks such as that by Ojewale et al. (2025) to articulate intended outcomes for audit stages beyond evaluation; (2) design requirements that clearly specify actor responsibilities, minimum activities, and intermediary steps essential to achieving desired outcomes; and (3) co-design standards with stakeholders beyond intended users to encourage more robust definitions of desired outcomes that counterbalance usersâ desires for flexibility. 2. Background 2.1. Probabilistic Genotyping Software Systems Forensic DNA testing aims to determine whether a person of interest (e.g., the defendant) could be the source of some or all of the DNA recovered from a crime scene (Butler et al., 2024). Since 2009, the U.S. criminal legal system has increasingly relied on probabilistic genotyping software (PGS) to analyze complex DNA evidence (Moss, 2015; Commonwealth v. Foley, 38 A.3d 882, 888â90, (Pa. Super. Ct. 2012)).333Moss (2015) describes, âIn 2009, TrueAllele made its first appearance in a United States courtroom in the murder case of Commonwealth v. Foley.â. PGS are now widely used in criminal cases around the country. For example, in March of 2025, the creators of STRmixâone of the two leading PGS alongside TrueAlleleâannounced that STRmix is being used by 91 forensic labs across the U.S., by 29 more labs internationally, and in over 690,000 criminal cases worldwide (STRmix, Mar 11, 2025) Figure 1 depicts how forensic labs integrate PGS in casework. A laboratory analyst first prepares two DNA samples for input into the software: an evidentiary sample taken from the crime scene and a reference profile from a person of interest. The analyst then provides these samples as inputs to the software alongside subjective parameters, most notably the analystâs best estimate for the âapparent number of contributors.â The PGS uses these inputs to produce a likelihood ratio (LR). This final output is significantly impacted by the analystâs decision-making, such as their estimate for the number of contributors, as well as upstream laboratory equipment (Butler et al., 2024; Paoletti et al., 2005; Coble et al., 2015).444For more background on likelihood ratios and their use in forensic DNA statistics, see Section 2.5, âLikelihood Ratios: Introduction to Theory and Application,â in Butler et al. (2024). The LR compares the probability of observing the evidentiary sample under two competing hypotheses: one where the person of interest is a contributor and one where they are not.555For example, an LR may compare a hypothesis that the person of interest and two unknown individuals contributed to the DNA mixture, against a hypothesis that three unknown individuals who are not the person of interest contributed to the DNA mixture. LR values greater than 1 are interpreted as providing support for inclusion, while values less than 1 support exclusion. An LâR=1LR=1 is considered âuninformative.â We define a false negative result as an exclusionary LR (LâR<1LR<1) for a true contributor (i.e., a person of interest whose DNA is in the evidentiary sample) and a false positive as an inclusionary LR (LâR>1LR>1) for a non-contributor (i.e., a person of interest whose DNA is not in the evidentiary sample). 2.2. Validation of PGS Systems In forensics, validation is the empirical testing of a method prior to its deployment in casework. For PGS, this process consists of two parts: developmental validation (conducted by software developers) and internal validation (conducted by forensic laboratories seeking to use the software in casework) (Butler et al., 2024; Coble and Bright, 2019). In this work, we focus on internal validation studies, which we conceptualize as audits666In this paper, we use Birhane et al. (2024)âs definition of an audit as âany independent assessment of an identified audit target via an evaluation of articulated expectations with the implicit or explicit objective of accountability.â For a detailed explanation of the terms used in this definition, see Section I.A in Birhane et al. (2024).. 2.2.1. Purpose of internal validation studies A forensic lab must complete an internal validation study before using a PGS in casework. A goal of internal validation is to establish the operational rangeâthe range of DNA profile characteristics for which the lab has demonstrated acceptable performance (Butler et al., 2024). Dozens of profile characteristics are known to impact the reliability of PGS. Examples of those known to have the greatest impact include the number of contributors to the sample (NoC), the ratio of DNA amounts of contributors to the sample (if the sample is a DNA mixture, i.e., has more than one contributor), and the total amount of DNA in the sample (Butler et al., 2024).777For a list of notable profile characteristics, see Table 4.1 in Butler et al. (2024). Internal validation studies and the operational range established using study results can play a critical role in judgesâ decisions about whether or not the PGS outputs are sufficiently reliable to be presented as evidence at trial (Roth, 2016). This has been the case in several federal courts. For instance, the courts in U.S. v. Ortiz (United States v. Ortiz, 736 F.Supp.3d 895 (21-CR-2503), (S.D. Cal. 2024)) and U.S. v. Johnston (United States v. Johnston, (23-CR-13), (E.D. N.Y. 2025)) excluded PGS outputs due to sample characteristics in these cases exceeding the operational bounds established by the forensic laboratoriesâ internal validation studies. Likewise, the court in U.S. v. Lewis (United States v. Lewis, 442 F. Supp. 3d 1122 (18-CR-194), (D. Minn. 2020)) admitted some evidence and excluded some evidence after similar considerations. An internal validation studyâs compliance with an established standard can provide additional assurance to a judge that the forensic lab performed a high quality internal validation study and that the operational range is reliable. 2.3. Standards and Guidelines for PGS Validation In 2009, the National Academy of Sciences issued a report identifying a wide range of shortcomings in U.S. forensic science research and practice (Council et al., 2009). Recommendations six to eight of this report emphasize the need to establish forensic science standards, to accredit laboratories adhering to those standards, and to establish quality assurance and quality control practices for individual practitioners in order to promote rigor and reliability in forensic science methods. Following the NAS report, standards have been developed for a range of forensic disciplines. However, recent work raises the concern that these efforts have not yielded meaningful improvements and instead threaten to maintain the status quo (Sinha, 2022; Morrison et al., 2020). We build on this work by exploring how standards for internal validation of probabilistic genotyping software may contain similar deficiencies. While several national and international groups have issued guidance for PGS validation (28; M. D. Coble, J. Buckleton, J. M. Butler, T. Egeland, R. Fimmers, P. Gill, L. GusmĂŁo, B. Guttman, M. Krawczak, N. Morling, et al. (2016); 69; 27; 59), unlike ASB 018, these standards and guidelines are not approved by or registered with the American National Standards Institute (ANSI). 2.3.1. The AAFS Standards Board (ASB) In 2015, the American Academy of Forensic Sciences established the AAFS Standards Board (ASB), a standards development organization tasked with developing standards for forensics practice in the U.S.888For more background on the creation of the ASB, see Appendix A.1. Importantly, as opposed to the guidelines referenced in the previous paragraph, ASB standards are approved by and registered with the ANSI. This means that the ASB develops standards following ANSI requirements for âopenness, balance, lack of dominance, due process, and consensusâ ((ANSI), ). These requirements seek to ensure that the standards development process is âequitable, accessible, and responsive to the requirements of various stakeholders,â and that standards developed by the ASB allow organizations to âmaintain autonomy from vested interest groupsâ (of Forensic Sciences (AAFS), ). 2.3.2. ASB Standard 018 In 2020, the ASB published the first edition of ASB 018, the Standard for Validation of Probabilistic Genotyping Systems999ASB 018 is publicly available for download at: https://w.aafs.org/sites/default/files/media/documents/018_Std_e1.pdf, which outlines a set of requirements for developmental and internal validations of PGS (ANSI/ASB Standard 018, 1st Ed., 2020). ASB 018 can influence internal validation study practices in several ways. On one hand, it can provide guidance to labs conducting internal validation studies. On the other hand, it can shape internal validation studies through external scrutiny that pressures labs to ensure internal validation compliance with the standard. The latter currently happens primarily through admissibility hearings in criminal cases, where a judge may interpret a labâs claim of compliance with ASB 018 as providing assurance of validation study quality.101010When PGS is introduced in a criminal case in the U.S., the court may hold an admissibility hearing to determine whether the evidence is sufficiently reliable to be presented in front of the jury. All federal courts and some state courts use the Daubert Standard to conduct this assessment, and one factor considered under the Daubert Standard is âthe existence and maintenance of standards controlling its operation.â (at Cornell University, 2023) The U.S. v. Anderson court conducted a Daubert hearing for one PGS software, and concluded: âthe Government has established that [the PGS tool] complies with the relevant standards issued by [SWGDAM] and ANSI/ASB. Accordingly, this factor weighs in favor of admissibilityâ (United States v. Anderson, 673 F. Supp. 3d 671 (21-CR-204), (M.D. Pa. 2023)). In this paper, we focus ASB 018âs effectiveness as a tool for assessing and providing assurance of internal validation study quality. Our investigation of ASB 018âs effectiveness in providing appropriate assurance about compliant studies builds on critical concerns raised by recent work highlighting that judges frequently hold uncritical perceptions of forensic software (Jin and Salehi, 2024) and work raising concerns that standards create âa danger [âŚ] that a court may not look further than the fact that a standard exists, and be misled into believing that conformity to a vacuous standard is indicative of scientific validity, even though it is notâ (Morrison et al., 2020, p.207). 3. Related Work 3.1. Audits and Accountability Researchers, law and policymakers, and civil society groups increasingly call for audits as mechanisms for identifying and mitigating risks of algorithmic systems and holding developers and users accountable for the algorithms they develop and use (Raji et al., 2022; Wyden, 2023). Birhane et al. (2024) define an audit as âany independent assessment of an identified audit target via an evaluation of articulated expectations with the implicit or explicit objective of accountability.â Audits are not restricted to a specific method or type of audit target (Ojewale et al., 2025; Birhane et al., 2024; Goodman and Trehu, 2022). An audit may focus narrowly on a specific model or technical system (Buolamwini and Gebru, 2018), or more broadly on the sociotechnical system encompassing technical components and human users (Lam et al., 2023). Audits may assess computational notions of fairness, but may also focus on other criteria such as functionality and legality (Radiya-Dixit and Neff, 2023; Ojewale et al., 2025; Abebe et al., 2022). Additionally, audits may be conducted by a variety of actors, including an entity within the organization that built the examined algorithmic system, or those outside the organization. Most importantly, an audit does not end with a thorough performance analysis. Instead, evaluation of the audit target is situated within a broader accountability process in which the audited entity or organization faces consequences or is otherwise responsible for acting on the outcomes of the evaluation (Ojewale et al., 2025). However, audits do not always successfully create accountability (Latonero and Agarwal, 2021; Selbst, 2021; Goodman and Trehu, 2022; Raji et al., 2022). When an audit is poorly designed or executed, a claim that a system has been audited can, at best, be meaningless, and at worst, lead to âaudit washingâ where the audit provides false assurance by lending credibility to a dysfunctional or otherwise inadequate system (Goodman and Trehu, 2022). For example, recent studies investigating audits of AI hiring systems performed in accordance with New York Cityâs Local Law 144 (L144) discuss how audits might be able to comply with the lawâs audit requirements despite being incomplete assessments of algorithmic bias (Wright et al., 2024; Gerchick et al., 2025). Collectively, this body of work calls for careful attention to how audits are designed and conducted, and highlights how inadequate audits can hide and lend credibility to dysfunctional or inadequate AI systems. We build on this work by exploring how the design of an auditing standard can contribute to audit washing by lending credibility to audits that can be interpreted to comply with the standard despite failing to produce the outcomes envisioned by the standardâs creators. 3.2. Designing Effective Tools for Responsible AI A growing body of work in human-computer interaction and computer-supported cooperative work seeks to design tools such as software, frameworks, and checklists to help various stakeholders promote fairness, transparency, and accountability throughout the development, deployment, and use of AI systems (Deng et al., 2022; Lee and Singh, 2021; Madaio et al., 2020; Raji et al., 2020; Gebru et al., 2021; Mitchell et al., 2019). Recent work examines these tools in the context of AI auditing, highlighting tools that support various stages of the auditing process such as audit standards that outline best practices, and repositories that facilitate broader communication of audit results (Ojewale et al., 2025). Several groups have explicitly highlighted the importance of audit standards as tools for improving audit quality and consistency, thereby minimizing the risk of audit washing (Costanza-Chock et al., 2022; Raji et al., 2022; Goodman and Trehu, 2022; Ojewale et al., 2025; Lam et al., 2024). However, numerous studies have shown that the existence of an RAI tool, alone, does not guarantee its effectiveness. A tool must be adopted and used, often within organizational structures and cultures that can hinder a toolâs effectiveness in practice (Wong et al., 2023; Madaio et al., 2020; Rakova et al., 2021; Metcalf et al., 2021). Collectively, this body of work emphasizes the importance of examining how a tool is (or might be) used in practice and designing tools with real-world organizational cultures, settings, and incentives in mind. In parallel and in response to these studies, recent work has called for the human-computer interaction community to expand its work to evaluating the effectiveness, and not just the usability, of these tools (Berman et al., 2024). 3.2.1. Designing Effective Audit Standards Similarly, the existence of an audit standard, or an auditâs compliance with a standard, does not necessarily ensure accountability (Ojewale et al., 2025). In fact, poorly designed standards can lead to ineffective audits and can lend further credibility to inadequate audits. Standards defined by the audit target (Goodman and Trehu, 2022; Hoofnagle, 2016), and standards with overly broad or vague language (Goodman and Trehu, 2022; Costanza-Chock et al., 2022; Raji et al., 2022; Wright et al., 2024; Morrison et al., 2020), can fail to ensure adequate audits and can add further credibility to an inadequate system. For example, Wright et al. (2024) examine audit requirements defined in New York Cityâs Local Law 144 and discuss how the language grants companies substantial discretion over whether their technology falls within the scope of the lawâs audit requirements. This discretion can undermine accountability by allowing companies to decide that a technology is out of scope, even if the system is potentially biased and meaningfully impacts employment outcomes. To ensure that auditing standards lead to high quality audits and provide appropriate assurance about compliant audits, groups have studied the outcomes of requirements for audits in existing laws and synthesized lessons from other auditing domains such as finance and medicine (Raji et al., 2022; Wright et al., 2024; Groves et al., 2024; Lam et al., 2024; Gerchick et al., 2025). This body of work has collectively identified several key considerations for the design of effective standards. Some groups have called for standards that better account for their usersâ needs and work contexts (Ojewale et al., 2025; Schor et al., 2024). Others have discussed how the design of audit requirements must enable auditors to draw unambiguous conclusions about the audit target and enable audit stakeholders to draw unambiguous conclusions about a completed auditâs compliance with the standard. Drawing on their experiences conducting audits under L144, Lam et al. (2024) emphasize that requirements must be âverifiable or observable conditionsâ that âmust enable auditors to form an unambiguous opinion about whether a given criterion is satisfied.â Wright et al. (2024) further emphasizes the importance of unambiguously verifiable criteria through their discussion of how the discretion that the language in L144âs requirements grants to companies makes it difficult for outsiders to evaluate whether the company complies with the law, a determination that drives critical actions such as regulator demands that a company stop using a given technology. Taken together, this line of work highlights the importance of verifiable and observable requirements that support unambiguous conclusions (Lam et al., 2024; Wright et al., 2024; Raji et al., 2022), requirements that collectively support holistic evaluations (Lam et al., 2024), and context-specific requirements that are compatible with the practices and contexts of those using the standard (Schor et al., 2024; Ojewale et al., 2025). We seek to expand these efforts to build an evidence base for the design of auditing standards by studying aspects of the design of auditing standards that undermine its effectiveness. 4. Methods 4.1. Data Collection 4.1.1. ASB 018 standard To understand ASB 018âs goals, we drew on both the standard itself and the factsheet that the ASB provides to accompany the standard. 4.1.2. PGS audits (i.e., internal validation studies) To identify existing audit practices, we gathered PGS audit reports from the American Society of Crime Laboratory Directors Validation and Evaluation Repository111111The repository can be accessed at this link: https://w.ascld.org/validation-evaluation-repository/.. This database is, to our knowledge, the only publicly available centralized and voluntary repository for U.S. forensic labsâ PGS audits. From this repository, we found five PGS audit reports, all for the STRmix PGS: the NYC Office of the Chief Medical Examiner (OCME) Labâs internal validation of STRmix v2.4 (32) and STRmix v2.7 (34), the Colorado Bureau of Investigation (CBI) Labâs internal validation of STRmix v2.5 (33), the Palm Beach Sheriffâs Office (PBSO) Labâs internal validation of STRmix v2.6.2 (57), and the Maryland State Policeâs (MSP) internal validation of STRmix v2.9.1 (45).121212Because the creators of STRmix claim that the software is being used in 91 labs in the U.S. (STRmix, Mar 11, 2025), we expect there to be somewhere on the order of 100 audit reports. It is important to note that at the time of our study, this database only contained five. This dataset covers a range of software versions, state and city forensic labs of different sizes and geographies, and years (2016-2024).131313The first U.S. forensic lab internal validation of STRmix was completed in 2014 (Butler, 2024). In total, we analyzed over 280 pages of audit reports. 4.2. Assessing Study Compliance The first author first assessed each audit report against relevant ASB 018 requirements to understand whether the audit could be interpreted as complying with the standard. Importantly, only the MSP labâs audit report (completed in 2024) (45) claims compliance with ASB 018. Three audits were completed before the publication of ASB 018 in 2020 (OCME v2.4 (32), CBI v2.5 (33), PBSO v2.6.2 (57)). While the OCME labâs v2.7 (34) audit report was completed in 2021, it does not mention ASB 018. Thus, our focus is on interpreting whether labsâ practices could be interpreted as complying with the standard. Furthermore, of the three audit reports completed before 2020, none of the labs uploaded updated reports to the repository after 2020, despite ASB 018 advising labs to review their audits and supplementing where necessary (ANSI/ASB Standard 018, 1st Ed., 2020). This not only suggests that labs with pre-2020 audits believe their audits to be compliant with ASB 018, but also that ASB 018 did not change labsâ existing practices. We further expand on this observation in Section 6.1. 4.2.1. Identifying line-level requirements of ASB 018 Adopting an approach similar to Lawrence et al. (2023), we extracted all of the individual line-level requirements in ASB 018. We interpreted any action proceeded by âshallâ as a line-level requirement, and focused only on requirements for internal validation studies. Next, we wrote out, as unambiguously as possible, how we interpreted each requirement. This step was crucial due to multiple ambiguities in ASB 018âs language. In general, when we identified ambiguity in a requirement, we used the most permissive interpretation. For example, to assess compliance with ASB 018âs requirement that labs âshall evaluate both the appropriate sample types [âŚ] and the number of samples within each type [and] shall base this evaluation on the intended application of the softwareâ (R4.1.6), we asked: âDoes any part of the report claim that the studyâs evaluated samples are appropriate for the intended application of the software?â We provide additional details and the full list of our interpretations of each of the line-level requirements used in our study in Appendix B.1. 4.2.2. Assessing study compliance with each of the requirements of ASB 018 For each audit, the first author read through the audit report, and resolved any confusions about report language and study design with the fourth author. Then, for each study, the first author identified a portion of the report that could be interpreted as fulfilling our interpretation of each of ASB 018âs line-level requirements. Throughout this process, ASB 018âs design also created challenges for comprehensively assessing compliance. For example, ASB 018 requires that samples studied ârepresent (in terms of number of contributors, mixture ratios, and the total DNA [amount]), the range of actual casework samples intended for analysis with the system at the laboratoryâ (ANSI/ASB Standard 018, 1st Ed., 2020). While this requirement is not subject to the ambiguities discussed in Section 4.2.1, it was difficult for us to comprehensively assess compliance with this requirement because none of the labs explicitly stated their labâs intended range of analysis. Since ASB 018 does not require that labs explicitly state the range of samples intended for analysis, we did not interpret reportsâ lack of specified ranges as non-compliance. Instead, following our permissive approach to assessing study compliance, we accepted any language in the report that suggests the lab chose samples representative of their intended range (e.g., report language concluding that the audit demonstrated the software was âfit for its intended purpose in the labâ) as evidence of study compliance with this requirement. Table 1 in Appendix B.2 documents sections of each audit report that we took to fulfill each of our interpretations of ASB 018âs line-level requirements. 4.3. Data Analysis Our study investigates the gaps between internal validation practices that ASB 018 envisions and the internal validation practices it enables, and how the design of ASB 018 enables these gaps. To do so, we first identified ASB 018âs goals for compliant audits, i.e., what practices ASB 018 envisions or seeks to ensure about compliant studies. Next, we examined how labs have conducted audits in practice. Last, we returned to the standard and examined how its design enables the gaps between the envisioned and realized practices identified in the previous two steps. We analyzed ASB 018 using five audit stages from Ojewale et al. (2025)âs framework for stages of an audit: Audit Scope, Standards Identification, Performance Analysis, Audit Judgment, and Audit Communication. We used these stages to connect our study to existing literature and ensure our findings are relevant to other audit contexts, and we focused on this specific subset of stages given their relevance to the PGS context we are studying. For each audit stage, the first author identified and qualitatively coded relevant excerpts from the standard and the factsheet. We also coded these documents to investigate how ASB 018 envisions the purpose of audits. The first author then identified relationships between codes within each audit stage (and also for audit purpose) and grouped codes together into increasingly abstract themes through an iterative, bottom-up affinity diagramming process (Beyer and Holtzblatt, 1999). The group then further refined these themes through discussion, focusing on themes that captured key meanings within each audit stage and assessing the fit of each theme with its corresponding audit stage. We used our themes for ASB 018âs vision for audit purpose to complement our construction of themes for the standardâs goals for each audit stage. For example, we drew on our interpretation of the standardâs emphasis on the role of audits in âestablishing boundaries on software useâ in our interpretation of one of the standardâs goals as âexpectations are falsifiableâ. Next, to explore potential gaps between laboratoriesâ validation practices and ASB 018âs goals, the first author identified audit excerpts relevant to each of the ASB 018 themes and inductively coded those excerpts. Instead of looking for specific practices, we grounded our interpretations in the practices observable in the report. For each audit stage, the team discussed the initial list of codes to ensure codes captured meaningful divergences from the standardâs goals and then iteratively grouped codes into higher-level themes. Finally, we sought to understand how the design of the standard enables the gaps we identified in the previous step. For each ASB 018 theme, the first author identified relevant ASB 018 requirements. The first author analyzed these excerpts alongside the codes and high-level themes describing internal validation practices, and generated codes capturing specific characteristics of the requirements that enable observed patterns in validation study practices. The first author then iteratively grouped codes into themes. For example, under the Audit Scope stage, a theme representing ASB 018âs goals was âtest broader sociotechnical systemâ (corresponding codes included âinternal validation as providing assurance about systemâ and âacknowledging impact of user inputâ), a theme representing internal validation practices was âtesting software only, not including userâ (codes included âusing ground truth number of contributor valueâ), and a theme representing aspects of ASB 018âs design that enabled the mismatch between goal and practice was âisolated requirementsâ (codes included âaddressing performance measurement and user input separatelyâ). 4.4. Study Limitations and Future Work 4.4.1. Audit artifacts as proxies for audit practice We seek to understand current audit practices by examining how they are discussed and represented in publicly available audit reports. In doing so, our findings our based on accounts of practices. Because these reports are typically submitted in court, we expect them to accurately represent auditing practices. We hope that our findings can compel increased transparency about audit practices so future work can examine audit practices using a variety of sources and methods to investigate the extent to which audit reports authentically and fully capture audit practices. 4.4.2. Representativeness We analyzed audit reports produced by labs of varying types (city, county, and state), geographies, and sizes (e.g. large and small labs), for a variety of versions of the STRmix probabilistic genotyping software. We observed many similarities in the language and structure of the audit reports we examined, and we attribute these similarities to labsâ (i) reliance on SWGDAM validation guidelines, a widely used alternative source of guidance for internal validation studies, and (i) use of a reporting template provided by STRmix developers. These observations suggest that our findings may extend to audits that also reference the SWGDAM guidelines and audits by labs that also use the STRmix software. However, we caution against assuming that these findings are applicable to all validation studies. At present, while the creators of STRmix claim that it is being used in 91 labs in the US and 29 more internationally (STRmix, Mar 11, 2025), few labs have made their audit reports publicly available, which limits the extent to which we can conclude how representative these reports are (Butler et al., 2024). We hope that our findings can support efforts underway that are advocating for greater transparency of audit reports (e.g., (Canellas, 2021; Butler et al., 2024; Abebe et al., 2022; Jin and Salehi, 2024; Krane and Philpott, 2022; Matthews et al., 2019, 2020; Services, )) that would enable future work that can additionally leverage quantitative methods to explore the representativeness of our findings of the broader landscape of audit reports. 4.4.3. Our permissive approach to assessing study compliance While we drew on our collective familiarity and expertise in PGS and validation studies, our view does not necessarily represent the view of those who might actually be tasked with assessing laboratory audits for compliance with ASB 018, such as organizations and auditors tasked with forensic laboratory oversight and accreditation. Our study takes a permissive approach, towards our goal of identifying what practices are possible under ASB 018. This allows us to investigate the largest possible gaps between auditing practices and the standardâs desired practices, to raise awareness of what could possibly happen and to help safeguard against these possibilities. Future work can investigate how actors tasked with forensic laboratory oversight and accreditation interpret ASB 018âs requirements and the extent to which their interpretations vary. 4.4.4. Critically examining ASB 018âs goals Our study aims to assess current audit practices against ASB 018âs goals for PGS audits, towards our goal of assessing ASB 018âs effectiveness and how the design of the standard may undermine its effectiveness at ensuring compliant audits meet its goals. A critical area for future work to investigate, to complement our assessment of the ASB 018âs effectiveness, is the degree to which ASB 018âs own goals comprehensively assess PGS system effectiveness. Figure 2. Summary of the gap between each ASB 018 goal (blue text starting with G:) and the corresponding compliant but unsatisfactory audit practices (red text starting with P:). We identify gaps between ASB 018âs envisioned practices and real-world audit practices for each of five stages of the audit process, which we draw from Ojewale et al. (2025). . Flowchart with five stages flowing from left to right: Audit Scope, Standards Identification, Performance Analysis, Post-Audit Judgment, and Audit Communication. For each stage, we frame the audit stage as a question that auditors would ask at the audit stage, and then we summarize our findings on the gaps between ASB 018âs goals and compliant but unsatisfactory audit practices. Each gap corresponds to a subsubsection of the Results. Question for Audit Scope is: âWhat is the audit target and what issues is the audit prioritizing?â The two Goal-Practice gaps under Audit Scope are (1) Goal: Sociotechnical system, Practice: Inadequate evaluation of software user; (2) Goal: Samples representative of casework samples, Practice: Convenience Datasets. Question for Standards Identification is: âWhat expectations is the audit target evaluated against?â The three Goal-Practice gaps under Standards Identification are (1) Goal: Define expectations for system performance, Practice: Few explicitly stated expectations; (2) Goal: Expectations match lab customersâ expectations, Practice: Labs and their customers may have conflicting expectations; and (3) Goal: Falsifiable expectations, Practice: Subjective, contradictory expectations. Question for Performance Analysis is: âWhat data about system performance does the audit produce?â The two Goal-Practice gaps under Performance Analysis are (1) Goal: Measure performance on a variety of sample types, Practice: Evaluate a narrow subset of possible sample types; and (2) Goal: Sufficient sample size to address variability, Practice: No discussion of sufficiency. Question for Post-Audit Judgment is: âHow are evaluation results translated to judgments, consequences, or actions?â The two Goal-Practice Gaps under Post-Audit Judgment are: (1) Goal: Assess whether measured behavior is (un)acceptable, Practice: Frame system failures as expected and acceptable; and (2) Goal: Establish boundaries on software use based on unacceptable performance, Practice: No such boundaries. Question for Audit Communication is âHow are audit data, methods, findings, and conclusions documented?â The one Goal-Practice gap under Audit Communication is (1) Goal: Supports scrutiny and reproducibility by third parties, Practice: Insufficient information. 4.5. Positionality All authors were either trained as researchers in or work in the United States. The first author is a researcher in algorithm auditing and human-computer interaction. The second author is a researcher trained in statistics, and the fifth author is a researcher in empirical machine learning who has also served as an AI and law advisor to numerous defense, civic, and government organizations. Collectively, these authorsâ background and familiarity with auditing literature informed their decision to analyze ASB 018 and audit reports through the lens of five stages of the audit process and shaped the patterns they identified in the standardâs design and the auditing practices. The third and fourth authors are experts in forensic DNA profiling and PGS, and the third author is also a researcher trained in molecular biology and population genetics: they have written numerous articles on the subject in forensics literature and frequently testify as expert witnesses in criminal cases involving evidence produced by forensic lab use of PGS. Throughout their casework, they have examined dozens of laboratory PGS validation studies and have encountered a variety of laboratory uses of PGS. Our specific focus on ASB 018 is driven by their experiences observing laboratory references to ASB 018 in criminal cases and their perspectives on ASB 018 as a standard with qualities that meaningfully differentiate the standard from other existing PGS validation guidelines in forensic DNA. Their expertise and experiences with investigating real-world uses of PGS in criminal cases also focused our analysis of audit practices on practices they commonly observed that created meaningful risks to defendants subject to evidence produced using these software. 5. Results: Gaps between ASB 018âs goals (envisioned audit practices) and the audit practices it enables From our analysis of study compliance with ASB 018, we found that all the audits we examined could be interpreted as compliant with our permissive interpretation of ASB 018âs requirements language. In this section, we present our findings on the gaps between ASB 018âs envisioned audit practices and the audit practices it enables, and how the design of ASB 018 enables these gaps. First, we present our findings on how ASB 018 envisions the purpose of audits, since we use this to inform our interpretation of ASB 018âs goals for each of the five audit stages. We organize the remainder our findings by the five audit stages outlined in Section 4. In each subsection, we first present each of the goal-practice gaps and then discuss aspects of ASB 018âs design that enable these gaps. Figure 2 summarizes our findings on the goal-practice gaps for each audit stage. Audit Purpose: Audits help laboratories (i) establish boundaries on software use that delineate the range of DNA samples for which the audit has demonstrated acceptable PGS system performance. These boundaries, in turn, (i) shape protocols for software use in casework and (i) provide assurance to stakeholders that the likelihood ratio evidence produced using the PGS system in an individual criminal case is reliable. ASB 018 explicitly states this goal when defining internal validation as the âacquisition of test data within the laboratory to [âŚ] determin[e] [âŚ] limitations of the systemâ (ANSI/ASB Standard 018, 1st Ed., 2020, p.2). The factsheet provides additional context suggesting that these limitations are boundaries on software use in casework: validation studies âdefine limitations for its useâ that âestablish the range of DNA profiles upon which the program may be used effectivelyâ (21). Through these statements, we understand limitations as delineations between the types of DNA profiles for which internal validation studies have demonstrated acceptable PGS system performance that others may âhave confidence inâ (21), and, the types of DNA profiles for which the software has not been shown to be sufficiently reliable and should therefore not be used. 5.1. Audit Scope: Determining the Audit Target and Selecting Issues to Prioritize Investigating in the Audit. 5.1.1. Gap: Broader sociotechnical system vs. Probabilistic genotyping software. Goal: Compliant audits evaluate the broader sociotechnical system that takes a crime scene sample as input and produces a likelihood ratio (LR) as output. This system encompasses both the PGS and the PGS user. When using the PGS to produce a LR, a forensic lab analyst specifies key inputs to the software such as their estimate for the number of contributors (NoC) to the DNA sample. ASB 018 highlights the âuser input parametersâ as key components of the broader sociotechnical system that audits must âevaluat[e]â (ANSI/ASB Standard 018, 1st Ed., 2020, R4.1.4) and âaddressâ (ANSI/ASB Standard 018, 1st Ed., 2020, Annex R4.1.4). Testing this broader sociotechnical system is also crucial to ASB 018âs end goal for audits to provide assurance about the reliability of LR evidenceâa system outputâpresented in a criminal case. Practices: Several audits did not use the userâs estimate of the number of contributors (NoC) parameter: a key user input that significantly impacts the reliability of the final LR output and is frequently estimated incorrectly. While the OCME and CBI labsâ audit reports describe using user estimates of the NoC when testing software sensitivity and specificity, the PBSO and MSP labsâ audit reports describe using the ground truth NoC. These auditsâ use of the ground truth NoC is especially concerning given the frequency of incorrect NoC estimates in the audits that used user NoC estimates and their consequent impacts on the LR outputs. For instance, the OCME v2.7 audit report reveals that users underestimated the NOC for 25 out of 40 five-person samples and noted that several false negative results âoccurred with [five-person DNA mixtures] where the apparent NoC was [underestimated]â (34, p.21). Some audits claimed that software users would be able to prevent observed software failures in casework, despite having no experiments that tested software usersâ ability to identify and prevent such failures. For example, the OCME v2.7 audit observed several false negative LRs but claimed that a âtrained analystâ would be able to notice that the softwareâs intermediate outputs were âunintuitive,â re-run the software with a correction, and produce correct results (i.e., true positive LRs) (34, p.8). However, without incorporating the analystâs inspection of these intermediate outputs into the testing process, the audit provides no assurance that a user would correctly identify and fix such failures in casework. In other words, the audit report claims that the system would produce a correct output when analyzing similar samples in casework but does not conduct tests that would provide evidence to support the claim and provide assurance of system reliability in casework. We also found evidence in several reports that suggested that the PGS developers, and not the labs, conducted parts or all of the audit. This suggests that any PGS user decisions that were used in the audit may have come from the developers or other company employees, as opposed to the labâs analysts who would be operating the software in casework. For instance, the CBI v2.5 audit report begins with a memo from the labâs quality director stating: â[ESR] submitted the summary and associated data files for the [audit] for the [laboratory]â (33, p.1). The companyâs copyright on the report further suggests developer involvement in the study. While the OCME v2.4 audit report does not explicitly state developer involvement, presence of the ESR copyright also suggests developer involvement (32). This evidence suggesting that developers may have conducted part or all of these audits raises concerns that even when tests incorporated user estimates, the users in the audit may not have been following the same protocols or training in the same environment as laboratory analysts who would be using the software in casework. 5.1.2. Gap: Inputs representative of casework samples vs. Inputs chosen out of convenience. Goal: Compliant audits test the PGS system using inputs that represent the range of samples the lab (i) plans to analyze in casework and (i) will likely encounter in casework. ASB 018 requires that audits âinclude case-type profilesâ (ANSI/ASB Standard 018, 1st Ed., 2020, R4.1.3), which the standard defines as samples âexhibiting features that are representative of a plausible range of casework conditions [such as] [âŚ] shared allelesâ (ANSI/ASB Standard 018, 1st Ed., 2020, p.1). ASB 018 also expects that these profiles ârepresent [âŚ] the range of actual casework samples intended for analysis with the system at the laboratoryâ (ANSI/ASB Standard 018, 1st Ed., 2020, R4.1.3). Practices: Audits primarily constructed test samples out of convenienceânot out of a formal assessment of samples the lab intends to analyze or is likely to encounter in casework. Labs frequently encounter mixtures containing DNA from related individuals in casework, and analysis of these mixtures is more likely to produce false positives due to individuals having similar DNA profiles (i.e., sharing alleles) (Kelly et al., 2022). While the CBI v2.5 and MSP v2.9.1 audits constructed mixtures specifically to test scenarios where contributors to the mixture are relatives (33; 45), the PBSO v2.6.2 audit did not explicitly test system performance on mixtures with related contributors and instead described the mixtures in their test set as having âvarying amounts of allele sharingâ (57, p.10). The OCME v2.4 and v2.7 audit reports use similar language (32; 34). This language suggest that these audits did not design their test set to account for mixtures of related individuals, especially since any set of mixtures is likely to have some amount of allele sharing (i.e., DNA profiles of any two unrelated individuals are likely to contain some of the same alleles). 5.1.3. How ASB 018âs design enables these gaps ASB 018âs narrow definition of the audit target misaligns with its vision that audits provide assurance of system reliability in casework. ASB 018 acknowledges the importance of incorporating the software userâs decisions into audits and implies a broad audit scope when describing audits as providing assurance of the reliability of the LR output in criminal cases. However, ASB 018 defines âprobabilistic genotyping systemâ as the âsoftware, or software and hardwareâ (ANSI/ASB Standard 018, 1st Ed., 2020, p.2), failing to acknowledge the lab analyst whose use of the PGS shapes the ultimate LR output. The organization and language of ASB 018âs requirements treat performance measurement and user inputs as separate items for audits to âinclude.â ASB 018 first requires that audits âaddress [âŚ] accuracy, sensitivity, specificity, and precisionâ (ANSI/ASB Standard 018, 1st Ed., 2020, R4.1.3). In a separate requirement, ASB 018 requires that audits also âinclude evaluati[ons of] user input parametersâ (ANSI/ASB Standard 018, 1st Ed., 2020, R4.1.4). Instead of requiring that labs account for and measure how user inputs impact system accuracy, sensitivity, specificity, and precision, ASB 018 treats these two components as two separate criteria that audits can address in isolation of one another. For instance, labs can satisfy the first requirement by measuring PGS performance without considering user inputs and the second requirement with a separate study that changes user input parameters and observes whether LRs increase or decrease without measuring performance (57; 45). ASB 018 uses vague language that allows labs to do very little to comply. As discussed earlier, ASB 018 requires that audits âshall include evaluating user input parameters,â but does not specify what this evaluation must entail (ANSI/ASB Standard 018, 1st Ed., 2020, R4.1.4). ASB 018âs requirement for labs to âconsider the effect of overestimating and underestimating the number of contributorsâ is similarly vague. There is no requirement to measure the impact of user inputs on system accuracy, sensitivity, specificity, and precision. ASB 018 envisions that labs conduct audits alone, but its use of vague language and failure to consider how developers might conduct part or all of an audit does not prohibit audits conducted by developers. ASB 018 frames most requirements at the level of the internal validation study without specifying the actor responsible for carrying out the required action. When ASB 018 does address the lab, it does not clearly specify which audit components the lab is responsible for. Take, for example, the requirement that âthe laboratory shall validate [the] PGS systemâ (ANSI/ASB Standard 018, 1st Ed., 2020, R4.1). Does a labâs review of a summary of the audit conducted by the developers count as the lab âvalidatingâ the software? Can the developers provide the inputs to the software using their estimates for user inputs such as the number of contributors? ASB 018 does not require that labs take note of the scenarios they typically encounter or intend to analyzeâa preliminary step that would hold labs accountable to some specific range of samples. ASB 018 requires that audits test âcase-type profiles [âŚ] that represent [âŚ] the range of actual casework samples intended for analysisâ (ANSI/ASB Standard 018, 1st Ed., 2020, R4.1.3), but does not require labs to determine and specify this rangeâa preliminary step that would help hold labs accountable for creating test sets that represent this range of intended uses. 5.2. Standards Identification: Articulating the Criteria or Expectations the System is Held to in the Audit 5.2.1. Gap: Predetermine specifications that define acceptable system performance vs. Failure to explicitly define specifications Goal: Labs test the PGS system against predetermined specifications: expectations for acceptable system performance. In its definition of âvalidation,â ASB 018 emphasizes the importance of ensuring the PGS system âwill consistently [meet] its predetermined specificationsâ (ANSI/ASB Standard 018, 1st Ed., 2020, p.2). Taken together with Requirement 4.1.3 and ASB 018âs vision for audit purpose, we understand ASB 018 to further suggest that these specifications must define acceptable âaccuracy, sensitivity, specificity, and precisionâ of the PGS systemâs LR output (ANSI/ASB Standard 018, 1st Ed., 2020, R4.1.3). Practices: Audit reports often do not explicitly define expectations for acceptable software behavior, suggesting that audits were not evaluating the system against predetermined expectations for acceptable performance. Instead of explicitly stating the expectations they were evaluating against, audits sometimes implicitly communicated their expectations by describing observed results as âexpected.â For example, after presenting observed true positive rates, the MSP v2.9.1 audit report concludes: âAs expected, there was a high percentage of [true positive] LRsâ (45, p.27). This language suggests that the lab may have understood a âhighâ true positive rate to be acceptable system sensitivity, but the audit never explicitly stated this expectation. At other times, audits simply concluded by presenting results without any language suggesting that results met expectations, making it difficult to infer what the labâs expectations were or whether they had any expectations to begin with. For example, the OCME v2.4 audit measured software precision (i.e., the degree of variability in LRs between multiple software runs on the same inputs) and concluded that the observed LRs âdemonstrate that [âŚ] the values obtained for the various runs remain close; within one order of magnitude except for the four-person mixtureâ (32, p.19). 5.2.2. Gap: Expectations align with forensic labsâ customersâ expectations for PGS performance vs. Expectations that may conflict with customersâ desires Goal: Compliant audits test the system against expectations for acceptable system performance that align with lab customersâ perceptions of what constitutes acceptable system performance. ASB 018 emphasizes that â[audits] provide the study results and conclusions necessary for customers of [forensic labs] to have confidence in the [LR] evidence providedâ (ANSI/ASB Standard 018, 1st Ed., 2020, Foreword). In order for audits to support customer141414We interpret ASB 018âs use of the word âcustomerâ to imply stakeholders who pay laboratories for their services (e.g., prosecutors). Importantly, ASB 018 does not use the word âstakeholderâ, which refers to a broader set of actors. assurance of PGS system reliability, the audit must assess PGS system behavior against lab customersâ expectations (as opposed to the PGS developer or forensic labâs expectations). Practices: Audit reports suggest that audits relied on expectations for acceptable system performance that may misalign with customersâ needs, and none of the audit reports mention customersâ needs or perceptions. For example, the MSP v2.9.1 audit report concludes that the auditâs observed false positive rates of up to 1.3% and false negative rates of up to 16% demonstrate that the PGS performed âas expectedâ (45, p.27). In doing so, the audit did not consider how a defense attorneyâs or a prosecutorâs perception of acceptable false positive or false negative rates may differ. Furthermore, we found that none of the audits even mentioned customersâ needs or perceptions in their reports, suggesting that labs did not consider whether their expectations aligned with those of their customers. 5.2.3. Gap: Falsifiable expectations vs. Subjective, otherwise unclear, and contradictory expectations. Goal: Expectations for acceptable system performance are falsifiable. In other words, labs define what constitutes acceptable system behavior using criteria that enable labs to assess observed system behavior against these criteria and unambiguously state whether the system meets these criteria. Crucially, labs must be able to unambiguously identify when software behavior fails to meet these criteria for acceptable performance, since the end goal of an audit is to establish boundaries on acceptable software use. Practices: Audits described expectations using ambiguous criteria. For instance, audits frequently defined expectations using subjective criteria such as âhighâ and âlow.â For example, all audits stated that the LR for true contributors âshould be high.â The MSP v2.9.1 audit report concludes that their observed sensitivity rates demonstrated âhigh [PGS] sensitivityâ (45, p.23). Audit reports also use other subjective criteria when discussing expectations, such as âintuitiveâ (e.g., âresult in an intuitive inclusionary LRâ (34, p.7)), âgenerallyâ (e.g., âtrue contributors generally gave high LRsâ (34, p.21)), and âcloseâ (e.g., âthe values obtained for the various runs remain closeâ (32, p.19)). Audits also used otherwise unclear language in their expectations. All auditsâ definitions of sensitivity imply that they expected the PGS to âreliably resolve the DNA profile of [true] contributorsâ (32; 34; 57; 45; 33). However, none of the audit reports specify what it means for a PGS to âresolveâ a DNA profile. For example, does an LR above 1 count as âresolvingâ the profile, or should the LR be above some other threshold value like 1,000? Expectations within the same report were sometimes contradictory. For example, the OCME v2.4 audit report defines specificity as âthe ability of the software to reliably exclude non-contributorsâ (32, p.6). However, when discussing experiment results, the report states that some false positive LRs were âexpectedâ (32, p.14). 5.2.4. How ASB 018âs design enables these gaps While ASB 018 refers to âpredetermined specificationsâ in its definition of validation, the standard does not include any requirement that labs actually define these specifications. Furthermore, despite emphasizing the importance of audits establishing boundaries on acceptable software use, ASB 018 does not establish any requirements for how these specifications should be defined to ensure that defined expectations actually support the end goal to establish boundaries on acceptable software use. Instead, ASB 018 simply requires that labs âaddress [âŚ] accuracy, sensitivity, specificity, and precisionâ (ANSI/ASB Standard 018, 1st Ed., 2020, emph. added). Similar to our discussion of ASB 018âs use of the word âconsiderâ in Section 5.1.3, the word âaddressâ allows labs to comply without clearly and explicitly defining what constitutes acceptable accuracy, sensitivity, specificity, and precision and then assessing PGS system performance against these predetermined specifications. 5.3. Performance Analysis: The Actual Evaluation Itself (Gathering Data on System Performance) 5.3.1. Gap: Measure system performance on a variety of sample types vs. Measure performance on a narrow subset of samples tested. Goal: Labs measure sensitivity, specificity, precision of the PGS systemâs LR outputs for a variety of sample types, and labs measure accuracy for a smaller set of sample types. A goal that follows from defining expectations for acceptable PGS system performance with respect to sensitivity, specificity, precision, and accuracy is measuring these dimensions of system performance. Furthermore, ASB 018 envisions that labs generate these measurements for a variety of DNA sample typesâa detail crucial to ensuring that audits help establish boundaries on software use. ASB 018 also envisions that labs measure system accuracy for a variety of sample types but explicitly limits its expectations for labs to a smaller set of sample types: â[one-person] samples [and simple][two]-person mixturesâ (ANSI/ASB Standard 018, 1st Ed., 2020, p.1). ASB 018 emphasizes that situations where âthe ground truth is not known are not suitable for accuracy studiesâ (ANSI/ASB Standard 018, 1st Ed., 2020, p.1) Practices: Audits often measured LR sensitivity, specificity, precision, and accuracy on a narrow subset of the samples that labs had access to or could have tested. Every report documents audit measurements of software sensitivity and specificity in an experiment specifically dedicated to sensitivity and specificity. In these experiments, most audits did not explicitly test DNA sample types that labs often encounter in casework, such as significantly degraded151515E.g., DNA samples degrade after exposure to sunlight for several days mixtures. Instead, reports only address degraded samples when describing later experiments that often did not measure system specificity, precision, and accuracy. For example, the OCME v2.4 audit artificially degraded a single one-person DNA sample and only investigated the impact of degradation on the true contributorâs LR (instead of also investigating non-contributor LRs to test specificity) (32). All audits investigated system precision in separate experiments, and most auditsâ precision experiment test sets were significantly smaller subsets of their sensitivity and specificity experiment test sets. Lastly, while ASB 018 explicitly constrains the set of sample types for which labs should measure system accuracy, ASB 018 does include simple two-person mixtures in its set of suitable sample types. However, all auditsâ accuracy experiments only evaluated one-person samples. Audits sometimes even acknowledged that the samples they were testing were a subset of the entire set of sample types they could have tested, stating something like: âThere is a small subset of profiles where [system accuracy can be tested]. These include [one-person samples]â (32, p.3). In other words, these audits measured accuracy on one type of sample (which is also one of the types of DNA samples least likely to be associated with poor PGS performance) when they could have also included two-person mixtures. 5.3.2. Gap: Robust audit measurements vs. Inadequate consideration of sample size sufficiency Goal: Labs produce robust measurements of system performance through repeated testing to account for variability in measurements. ASB 018 recognizes that performance measurements may vary due to upstream DNA preparation procedures, stochasticity in the softwareâs underlying statistical models, and differences between DNA samples. As such, the standard highlights the importance of ârepeated testingâ on a âsufficientâ range of sample types and number of samples of each type (ANSI/ASB Standard 018, 1st Ed., 2020, Annex A, R4.1.3). Practices: Audits neither commented on the sufficiency of the sample sizes used to test system performance, nor provided measures used to determine whether sample sizes were sufficient. This lack of discussion is especially concerning since sample sizes varied drastically between audits. For example, when testing system precision, the MSP v2.9.1 audit used one four-person sample (45, p.65), whereas the OCME v2.7 audit used eight samples spanning two- to five- contributor mixtures (34, p.28). For sensitivity and specificity, while the MSP v2.9.1 audit tested 30 two-person and 17 five-person mixtures (45, p.14-15), the OCME v2.7 audit tested 169 two-person and 40 five-person mixtures (34, p.15). 5.3.3. How ASB 018âs design enables these gaps ASB 018âs vague language and use of separate line-level requirements to address interrelated concepts allows labs to analyze PGS performance on narrow subsets of test sets. While ASB 018 requires that internal validation studies to address âaccuracy, sensitivity, specificity, and precisionâ in one line-level requirement, another line-level requirement states, âThese studies shall include [samples] that represent [âŚ.] the range of actual casework samples intended for analysisâ (ANSI/ASB Standard 018, 1st Ed., 2020, R4.1.3). In this second requirement, ASB 018 does not specify whether â[t]hese studiesâ broadly refers to audits or narrowly refers to individual accuracy, sensitivity, specificity, and precision experiments within these audits (which the standard also refers to as âstudiesâ elsewhere, e.g., âaccuracy studiesâ (ANSI/ASB Standard 018, 1st Ed., 2020, p.1)). A permissive interpretation of the second requirement would interpret â[t]hese studiesâ to mean â[i]nternal validation studies,â which would allow for a precision experiment to test a narrow subset of test samples while the audit as a whole satisfies the requirement for audits to test representative DNA samples. Furthermore, as we have discussed earlier (Section 5.1.3), ASB 018 treats complex sample characteristics such as degradation and software performance measurement as two separate criteria that audits can address in isolation from one another. ASB 018 lists only three dimensions of DNA sample types in the requirements, failing to capture key dimensions that the standard outlines in its Terms and Definitions section. The standard requires that audits include case-type profiles âthat represent (in terms of number of contributors, mixture ratios, and total DNA [amount]) the range of actual casework samples intended for analysisâ (ANSI/ASB Standard 018, 1st Ed., 2020, R4.1.3, emph. added). Here, the standard names three DNA sample characteristics that labs must account for: the number of contributors, mixture ratios, and the total DNA amount. However, the standardâs definition of âcase-type profilesâ names other dimensions to consider, such as the amount of âdegradationâ and the number of âshared allelesâ between individual contributors in a mixture (ANSI/ASB Standard 018, 1st Ed., 2020, p.1). ASB 018âs omission of key factors like degradation and allele sharing not only limits attention to the dimensions specified in the requirements, but also introduces ambiguity around whether the dimensions mentioned in the Terms and Definitions section are requirements or suggestions. ASB 018 neither provides criteria for unambiguously determining whether the sample size of a labâs test set is sufficient, nor requires labs to establish their own criteria. Instead of specifying a minimum number that labs must comply with, ASB 018 requires that labs âshall perform sufficient studiesâ (ANSI/ASB Standard 018, 1st Ed., 2020, Annex A, R4.1.3). However, the standard neither specifies how labs should operationalize âsufficient,â nor requires labs to establish their own criteria for determining sufficiency. While the standardâs requirements language provides labs with the flexibility to design studies that fit their needs and contexts, it also enables labs to test any types and number of samples (e.g., two samples) without justifying why their choice is sufficient. In other words, the standard does not hold labs accountable to any definition of âsufficient.â 5.4. Post-Audit Judgment: Discussing Results and Translating Results Into Actions or Consequences 5.4.1. Gap: Conclude whether measured behavior is acceptable vs. Frame software failures as correct behavior. Goal: Labs assess observed behavior and determine whether observed behavior is acceptable or unacceptable. Providing judgment about observed system behaviors, rather than stopping at simply producing these observations (i.e., measuring system behaviors), is crucial to ensuring that audits establish boundaries on acceptable software use in casework. In other words, the lab must label a test result as acceptable or unacceptable, instead of simply stating the test result. Practices: Audits frequently framed software failures as correct behavior. Most audits did not use the word âfailureâ to describe incorrect results, false positives, and false negatives. For instance, the OCME v2.4 audit report frequently reframes observed false positive results as correct behavior. The audit report states that the false positives were ânot a failure of STRmixâ (32, p. 13). The OCME v2.7 audit report uses similar language (34, p. 24) and also frames other software failures as ânon-intuitive,â stating that some test runs âresulted in non-intuitive results with an LR = 0 or failure to runâ (34, p.14). Audits also frequently framed software failures as ânot unexpectedâ instead of as failures, responding to false positive and false negative LRs with statements like âthese results are not unexpectedâ (34, p.24). None of the audit reports analyzed in our study frame observed failures as âunacceptable.â 5.4.2. Gap: Establish boundaries on acceptable software use in casework vs. Not establishing such boundaries Goal: Labs use their judgments of observed system behavior to establish boundaries on acceptable software use in casework. As we described at the beginning of Section 5, ASB 018 envisions that audits ultimately establish boundaries on acceptable software use, and these boundaries must be specified in terms of DNA sample types. Practices: None of the audits used observed software failures to define boundaries on software use in terms of DNA sample types. Some audits established another type of boundary: an âuninformativeâ range for LRs (e.g., 0.001 - 1000 (34, p.24)) that the lab would use in criminal cases to present LRs falling in this range as âinconclusiveâ, i.e., neither inculpatory nor exculpatory. However, this boundary is not a restriction on software use and does not prevent labs from introducing false positive LRs as evidence in criminal cases when the PGS system produces extreme false positive LRs (which is common in the case of marginal samples where there is a high degree of allele sharing between a true contributor and a non-contributor). Oftentimes, in place of boundaries on software use, audits emphasized that analysts would be able to identify the issue and adjust their software use accordingly. For example, upon observing false positive results, one audit concluded that such scenarios are ârecognizable by trained analystsâ who would be able to correct the analysis accordingly to avoid such errors (34, p.8). However, as we discussed in Section 5.1.1, audits were scoped around the PGS instead of the broader sociotechnical system, so these audits do not actually provide assurance that analysts would be able to avoid such errors in casework. Audits also sometimes concluded without providing any assurance of reliability in casework, instead advising analysts to proceed with caution. For example, after observing several false negative results, the CBI v2.5 audit simply advised analysts to âexhibit caution [âŚ] for situations such as thisâ (33, p.31). 5.4.3. How ASB 018âs design enables these gaps ASB 018 neither clearly defines âspecificationâ, nor requires labs to explicitly label whether observed performance is acceptable. In Section 5.4.1, we discussed how labs framed software failures as ânot unexpected,â instead of labeling these failures as unacceptable behavior that the lab should prevent in casework through establishing boundaries on acceptable software use based on audit results. These practices further reveal consequences of the standardâs failure to specify what constitutes a âspecificationâ (see Section 5.2.4), and also reveal that the standard fails to ensure that labs actually judge whether audit results are acceptableâa crucial step towards establishing boundaries on acceptable software use. ASB 018 does not explicitly require that audits establish boundaries on software use. Instead, ASB 018 relies on the word âlimitationsâ, which the standard does not define. Despite repeatedly describing that audits help determine limitations, ASB 018 does not define what constitutes a limitation. The accompanying factsheet clarifies that limitations are boundaries that define âthe range of DNA profiles upon which the program may be used effectivelyâ (21). However, these details are crucially lacking from the standard, the sole document that studies would be assessed against for compliance. Furthermore, despite emphasizing that audits establish âlimitationsâ, ASB 018 does not require a lab to actually specify the limitations they identify in the audit. In other words, the standard seeks to guarantee that labs use audits to determine system limitations, but makes no requirement that labs actually do so. 5.5. Audit Communication: Documenting Audit Design, Methodology, and Results 5.5.1. Gap: Documentation enables third party scrutiny and reproducibility vs. Documentation contains insufficient information about test samples and results Goal: Lab documentation of audit data, methods, and results enables third party scrutiny. ASB 018 emphasizes the importance of audit documentation that addresses the information needs of third parties, stating: âit is incumbent upon any laboratory performing these [audits] to retain these results for the examination and evaluation by third parties. The results should be documented in such a way that the [âŚ] [audits] can be reproduced and decisions made on the basis of these studies documentedâ (ANSI/ASB Standard 018, 1st Ed., 2020, p.5). Practices: Audit reports often do not provide sufficient information about test samples and conditions that third parties could use to interpret study results. The audit reports we analyzed omit crucial information about the range of DNA sample types examined, often describing test samples using vague language. For example, each audit report describes their sensitivity and specificity experiment test samples as âvaryingâ in their â[total] DNA [amounts,] mixture proportionsâ, and âamounts of allele sharingâ (e.g., (57, p.10)). When audit reports do provide specific ranges of sample types explored, they often only provide details for a small set of DNA sample characteristics (e.g., specifying the range of total DNA amounts studied, but not specifying the range of mixture proportions studied (34, p.15)). Audit reports also often do not provide data that would enable third party verification of their claims. For example, the OCME v2.4 audit report claims that âThe results of all comparisons was expected. [âŚ] the [non-contributors got] very low or 0 LRsâ (32, p.22). In the appendix, the audit report documents the number of non-contributors with LRs = 0 and plots the LRs in scatterplots, but does not provide exact LR values for non-contributors, despite the plots revealing several false positive LRs (32, p.47-50). 5.5.2. How ASB 018âs design enables these gaps ASB 018 requires that labs document their audits, but non-mandatory wording to describe the importance of ensuring that documentation is accessible and useful to third parties. ASB 018 requires documentation by stating that âAll [audits] shall be documented and retained by the laboratoryâ, but switches to using the phrase âit is incumbent uponâ when emphasizing that labs should document the results to support âexamination and evaluation by third partiesâ (ANSI/ASB Standard 018, 1st Ed., 2020, Annex A R4.5). ASB 018 also fails to specify specific pieces of information that labs must document, even though the standard specifies in the line immediately after that âLaboratories shall have a summary statement of the sample types of which the developer used to for their developmental validationâ (ANSI/ASB Standard 018, 1st Ed., 2020, Annex A R4.5) While the standard could specify a similar requirement that labs document the sample types they tested in their audit, the standard does not include any documentation requirement with this level of specificity. 6. Discussion Throughout our analysis of ASB 018 and five audit reports, we found that all five audits can be interpreted as compliant with ASB 018 despite falling short of the standardâs goals. We identified gaps between ASB 018âs envisioned practices and the practices it enables across five audit stages and highlighted how the design of the standard enables these gaps. First, we join Berman et al. (2024) in emphasizing the importance of tool developers providing clear definitions of the outcomes they intend for their tool to create. Our qualitative coding of ASB 018 and its accompanying factsheet to develop our interpretations of the goals of the toolâs creators was crucial to our evaluation of ASB 018 because the ASB did not explicitly state the standardâs goals. Without clear statements of intended outcomes, how can one conclude that a tool is effective? Furthermore, in situations like ours where tool developers do not clearly state goals, we contend that clearly articulating perceived goals from the outside, using methods such as our interpretive analysis of the toolâs goals, can help initiate conversations about desired outcomes and compel developers of audit tools to clarify their intended goals, towards holding tool creators (e.g., standards developers) accountable for the tools (e.g., standards) they create and promote. We additionally make the following recommendations, which we expand on in the sections below: ⢠Designers of audit standards and other forms of audit tooling should carefully consider how designing a tool for compatibility with usersâ current perspectives and practices can undermine the toolâs effectiveness. This is especially crucial for audit standards where ineffective standards can further legitimize existing practices. (Section 6.1) ⢠Evaluating an audit standardâs effectiveness against fine-grained articulations of the standardâs desired outcomes can help identify where increased specificity is crucial for achieving desired outcomes. Existing audit frameworks (e.g., (Ojewale et al., 2025; Raji et al., 2020)) can help increase the specificity with which we articulate these desired outcomes. (Section 6.2.1) ⢠Components of audit standards where increased specificity may be crucial to achieving desired outcomes are: (i) actors and actor responsibilities, (i) minimum activities, and (i) intermediary steps. (Section 6.2.2) ⢠Co-designing audit standards with groups beyond standards bodies and users of standards is crucial for counterbalancing usersâ desires for flexibility and can encourage more robust definitions of an audit standardâs goals. (Section 6.2.3) 6.1. Carefully consider how designing a tool for compatibility with tool usersâ current perspectives and practices can undermine the toolâs effectiveness. Prior work in human-computer interaction and algorithm auditing has called for the design of standards that more closely align with the needs and perspectives of those using the standard (Schor et al., 2024; Ojewale et al., 2025). For instance, Schor et al. (2024) highlight a gap between an algorithmic system transparency standard and system designersâ understandings of transparency. They call for standards bodies to more closely align standards with designersâ current practices and understandings of transparency. These calls echo recent observations by Berman et al. (2024) that much of existing RAI tool development and evaluation prioritizes tool usability, potentially at the cost of effectiveness. While designing standards to align with usersâ needs, workflows, and work contexts can be crucial for encouraging tool use and achieving intended outcomes (Madaio et al., 2020; Holstein et al., 2019; Rakova et al., 2021), our findings caution that designing a standard to enable flexibility and compatibility with usersâ current perspectives and practices can simultaneously undermine the standardâs effectiveness. While vague language supports user discretion and improves the standardâs compatibility with labsâ existing approaches, it does not hold labs accountable for carrying out critical auditing practices. For instance, while ASB 018âs requirement that labs âconsider the impact of over- and underestimation of the number of contributorsâ is compatible with labsâ existing practices of observing whether LRs increase or decrease when analysts over- and underestimate the number of contributors to the mixture, it fails to hold labs accountable for measuring how inaccurate estimations of the number of contributors impacts system accuracy, sensitivity, specificity, and precision (Section 5.1.3). We contend that identifying how a standardâs compatibility may undermine its effectiveness is especially critical in the context of standards for algorithm auditing, where ineffective standards not only may fail to improve practices, but at the same time may further legitimize existing practices. Three of the five audit reports we examined were completed before the publication of ASB 018 and one audit report published afterwards does not seem to engage with ASB 018, yet all four could be interpreted to comply with the standardâs requirements, suggesting that ASB 018 is compatible with existing lab practices: these labs would not need to significantly change their existing practices to comply with ASB 018âs requirements. In the context of ASB 018, this raises the question: Was the standard designed to improve practices or, as (Morrison et al., 2020) warn, was it designed to provide post hoc legitimacy to existing practices? In the next section, we discuss strategies that future efforts to design effective audit standards may leverage to balance flexibility and compatibility with effectiveness and accountability. 6.2. Recommendations for designing audit standards that effectively support accountability 6.2.1. Fine-grained articulation of the audit standardâs desired outcomes and evaluating the standardâs effectiveness at achieving those outcomes can help identify areas where increased specificity is crucial for achieving desired outcomes. Prior work on standards and algorithm auditing has framed increased specificity in standards as a double-edged sword: increased specificity can improve audit quality by mandating specific actions and outcomes, but can hinder auditor discretion that may be crucial to make context-appropriate, situational judgments (Raji et al., 2022; Costanza-Chock et al., 2022; Jackson and Barbrow, 2015). Based on our findings, we suggest that fine-grained articulation of outcomes desired by standards bodies and evaluating the standardâs effectiveness using these intended outcomes can help locate specific components where increased specificity is crucial for achieving the standardâs intended outcomes. For instance, whereas ASB 018âs goal for audits to establish boundaries on software use does not demand that labs use a specific ASB 018-mandated metric or threshold to define acceptable performance, it does require that labs specify some definition of what constitutes acceptable performance and that these specifications are falsifiable (Section 5.2.4) â forms of specificity that ASB 018 crucially omits. Similarly, whereas ASB 018âs goal for audits to test samples representative of the range of scenarios that the lab typically encounters or intends to analyze does not demand that all labs test up to five-person mixtures, it does require that labs specify a range that they and others could use to hold them accountable for. In the context of audit tooling, we recommend using existing auditing frameworks (e.g., (Ojewale et al., 2025; Raji et al., 2020; Abebe et al., 2022; Lam et al., 2024; Radiya-Dixit and Neff, 2023)) to define such goals. Using Ojewale et al. (2025)âs framework helped increase the specificity with which we articulated ASB 018âs goals and examined audit practices. For instance, instead of articulating only the desired outcomes with respect to audit outputs (e.g., establish boundaries), we were able to assess audit reports against practices the standardâs envisions for each audit stage (e.g., scoping the audit, defining performance expectations, and audit documentation). This also allowed us to address stages of the audit process beyond evaluation, such as standards identification and audit communication, that are crucial for creating accountability (Ojewale et al., 2025). Echoing calls by Ojewale et al. (2025) for tools that support all stages of the audit process, we argue that an audit standardâs goals should be articulated with sufficient granularity to address practices throughout the audit process. 6.2.2. Examine which actors, actor responsibilities, activities, and intermediary steps must be explicitly specified to accomplish the goals of the standard. Recent work studying algorithm audits in hiring has called for increased specificity in how rules for auditing (e.g., audit standards, requirements in audit legislation) define which algorithmic systems are considered in-scope (Wright et al., 2024; Groves et al., 2024) and who counts as an independent auditor (Groves et al., 2024). Building on this work, our findings suggest three components of audit standards where increased specificity may be crucial to ensuring audit standards produce standards developersâ desired outcomes and help create accountability: (i) actors and actor responsibilities, (i) bare minimum practices, and (i) intermediary steps. First, we found evidence that audit responsibilities were distributed between forensic labs and PGS developers in ways that undermined ASB 018âs goals (Section 5.1.1). Yet ASB 018âs requirements do not acknowledge and engage with this reality. Many of the standardâs requirements were directed at âinternal validation studiesâ (as opposed to specific actors), and the few requirements assigned explicitly to labs contained vague language that allows a wide range of lab involvement in audits. This lack of specificity with respect to actors and actor responsibilities not only undermines the effectiveness of the standard, but also risks papering over distributions of auditing responsibilities that raise additional concerns such as lack of auditor independence. Recent work by Groves et al. (2024) similarly calls for increased specificity with respect to various actors involved in auditing and their specific roles and responsibilities, highlighting that this specificity can help ease the relational work required of auditors who must navigate tensions between software developers, software users, and regulators. Based on our findings, we add that assigning explicit roles and responsibilities to specific actors can not only ensure desired audit practices, but can also make visible the various actors and their roles in shaping audits in ways that open auditing practices to assessment a wider group of stakeholders. Second, we found ASB 018âs use of vague requirements language such as âconsiderâ and âaddressâ consistently allowed for audit practices that fell short of the standardâs desired outcomes. For instance, the standard requires that labs âaddress [âŚ] precisionâ, and we see that labs analyzed a small subset of the test set used in the auditâs sensitivity and specificity experiment (Section 5.3.1). In fact, this vague language allows for even more cursory practices. For instance, in PBSOâs audit of STRmix v2.6.2, instead of measuring the precision of the software version being validated, the lab suggested that their previous validation of STRmix v2.4 and their sensitivity and specificity studies were sufficient. We recommend that designers of audit standards carefully assess the bare minimum an audit could do to fulfill a requirement, and reflect on what specific practices they seek to require instead of relying on vague language. Lastly, we found that labs omitted crucial steps necessary to achieve ASB 018âs goals, and ASB 018 enabled this through its failure to require those preliminary steps. For instance, ASB 018âs requirement that â[audits] shall include [samples] that represent [âŚ] the range of actual casework samples intended for analysis with the system at the laboratoryâ depends on labs first determining and specifying the range of samples they intend to analyze with the software, but ASB 018 does not require that labs do so (Section 5.1.3). These omissions also limited our ability to more thoroughly confirm study compliance (Section 4). For instance, similar to how Wright et al. (2024)âs ânull complianceâ signals a lack of publicly available information needed to make a concrete determination of compliance, we found that labs did not provide a specific range of DNA samples that we could use to concretely assess the range of DNA samples tested in the audit. Consequently, following our permissive approach to assessing compliance, we simply accepted a labâs claim that their test set covers the labâs intended range. Requiring these preliminary steps can ensure that those assessing audit compliance with a standard have crucial information needed to make unambiguous conclusions about audit compliance. To design audit standards that effectively produce desired outcomes and support rigorous, outside assessments of compliance, we recommend that standards developers consider breaking requirements out into smaller units to ensure that the standard adequately accounts for crucial preliminary steps that other requirements rely on. 6.2.3. Engaging stakeholders beyond those who will be using the standard can help complement usersâ perspectives and can uncover critical misunderstandings or disagreements over desired audit practices. As researchers and policymakers increasingly call for multistakeholder engagement in the design of audit standards (Schor et al., 2024; Hub, 2024) and RAI tooling more broadly, we argue that it is crucial that co-design efforts engage groups beyond the standardâs intended users. This can help ensure that usersâ desires for flexibility and compatibility with practice do not undermine effectiveness (as discussed earlier in Section 6.1). Furthermore, this broader engagement can help reveal how desired goals for audit practices may differ and conflict among individual members of standards-developing organizations (e.g., members of the ASB), users of standards (e.g., forensic laboratories), and various audit stakeholders (e.g., defense attorneys (Jin and Salehi, 2024; Abebe et al., 2022), prosecutors, independent experts in forensic DNA profiling (Krane and Philpott, 2022; Butler et al., 2024)). For instance, forensic laboratories may advocate for audit practices that make efficient use of laboratory resources, molecular biology researchers may call for more experiments on large sample sizes and diverse samples, and defense attorneys may advocate for more thorough documentation practices. In the context of the U.S. criminal legal system, building on work by Abebe et al. (2022) and Jin and Salehi (2024) that center the critical role of defense scrutiny of forensic software and call for adversarial audits (Abebe et al., 2022), we argue that it is especially crucial to engage defense attorneysâ adversarial perspectives, as adversarialism is crucial to protecting defendantsâ rights and the truth-seeking goals of the criminal legal system as a whole (Abebe et al., 2022). In other auditing contexts, engaging adversarial perspectives can not only help reveal standards developersâ blindspots and make visible a broader range of concerns, but can also highlight where there is less consensus in audit scope, methods, and goals (Geiger et al., 2024; Raji et al., 2022; Birhane et al., 2024). Frameworks that encourage fine-grained specification of desired outcomes can not only help standards bodies improve the specificity of their standards, towards creating effective audit standards, but can also help scaffold efforts to co-design standards with diverse stakeholders. Similar to how boundary objects for contested concepts like privacy (Mulligan et al., 2016) and fairness (Mulligan et al., 2019) can support communication and collaboration across diverse groups of stakeholders, we suggest that using frameworks for auditing (e.g., (Abebe et al., 2022; Ojewale et al., 2025; Raji et al., 2020)) can facilitate better understanding of potential misunderstandings or disagreements over core concepts, practices, and desired outcomes. 7. Conclusion This paper investigates how the design of an audit standard can undermine its effectiveness, complicate efforts to assess audit compliance with the standard, and undermine accountability through a case study of ASB 018, a standard for auditing probabilistic genotyping software. Through qualitative analysis of ASB 018 and five publicly available audit reports, we identify numerous gaps between the standardâs desired outcomes and the auditing practices it enables. We identify a number of design features that enable these gaps, such as the standardâs failure to define key concepts, requirements that treat audit components as separate considerations that can be assessed independently of one another, and vague language. Based on our findings, we recommend that designers of audit standards clearly articulate their goals, carefully examine how designing for compatibility with usersâ practices can undermine effectiveness, ensure sufficient specificity in the requirements language to ensure effectiveness and accountability, and co-design standards with stakeholders beyond the standardâs intended users. Acknowledgements.Thank you to Marc Canellas, Finale Doshi-Velez, Krzysztof Gajos, Jeanna Matthews, Andrea Roth, Rebecca Wexler, and the Harvard Human-Computer Interaction group for conversations that motivated and shaped this work. Thank you to Jennifer Friedman, Inioluwa Deborah Raji, Niloufar Salehi, and our anonymous reviewers for their comments and feedback on earlier versions of this paper. This material is based upon work supported by the National Science Foundation Graduate Research Fellowship Program under Grant No. DGE 2146752 and by the National Science Foundation under Grant No. 2243822. This research was also supported by the Andrew Carnegie Fellows Program, and the Hector Foundation. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the National Science Foundation or other funding agencies. References [1] A. N. S. I. (ANSI)American national standards (ans) introduction(Website) External Links: Link Cited by: §2.3.1. R. Abebe, M. Hardt, A. Jin, J. Miller, L. Schmidt, and R. Wexler (2022) Adversarial scrutiny of evidentiary statistical software. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, p. 1733â1746. Cited by: §1, §1, §3.1, §4.4.2, §6.2.1, §6.2.3, §6.2.3, §6.2.3. T. Ambrosius (Nov 4, 2025) External Links: Link Cited by: §A.1. ANSI/ASB Standard 018, 1st Ed. (2020) Standard for Validation of Probabilistic Genotyping Systems. Standard Vol. 1, AAFS Standards Board, Colorado Springs, CO. External Links: Link Cited by: 1st item, 10th item, 11st item, 12nd item, 13rd item, 14th item, 15th item, 16th item, 2nd item, 3rd item, 4th item, 5th item, 6th item, 7th item, 8th item, 9th item, §2.3.2, §4.2.2, §4.2, §5.1.1, §5.1.2, §5.1.3, §5.1.3, §5.1.3, §5.1.3, §5.1.3, §5.2.1, §5.2.2, §5.2.4, §5.3.1, §5.3.2, §5.3.3, §5.3.3, §5.3.3, §5.5.1, §5.5.2, §5, footnote 1. L. I. I. at Cornell University (2023) Daubert standard. External Links: Link Cited by: footnote 10. G. Berman, N. Goyal, and M. Madaio (2024) A scoping study of evaluation practices for responsible ai tools: steps towards effectiveness evaluations. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, p. 1â24. Cited by: §1, §3.2, §6.1, §6. H. Beyer and K. Holtzblatt (1999) Contextual design. interactions 6 (1), p. 32â42. Cited by: §4.3. A. Birhane, R. Steed, V. Ojewale, B. Vecchione, and I. D. Raji (2024) AI auditing: the broken bus on the road to ai accountability. In 2024 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), p. 612â643. Cited by: §3.1, §6.2.3, footnote 6. J. Buolamwini and T. Gebru (2018) Gender shades: intersectional accuracy disparities in commercial gender classification. In Conference on fairness, accountability and transparency, p. 77â91. Cited by: §3.1. J. Butler, H. Iyer, R. Press, M. Taylor, P. Vallone, and S. Willis (2024) DNA Mixture Interpretation: A NIST Scientific Foundation Review. NIST Interagency/Internal Report (NISTIR), National Institute of Standards and Technology, Gaithersburg, MD (en). External Links: Link, Document Cited by: §2.1, §2.1, §2.2.1, §2.2, §4.4.2, §6.2.3, footnote 4, footnote 7. J. Butler (2024) History of dna mixture interpretation: supplemental document to dna mixture interpretation: a nist scientific foundation review. Cited by: footnote 13. M. Canellas (2021) Defending ieee software standards in federal criminal court. Computer 54 (6), p. 14â23. Cited by: §4.4.2. A. Chouldechova, D. Benavides-Prado, O. Fialko, and R. Vaithianathan (2018) A case study of algorithm-assisted decision making in child maltreatment hotline screening decisions. In Conference on fairness, accountability and transparency, p. 134â148. Cited by: §1. M. D. Coble, J. Bright, J. S. Buckleton, and J. M. Curran (2015) Uncertainty in the number of contributors in the proposed new codis set. Forensic Science International: Genetics 19, p. 207â211. Cited by: §2.1. M. D. Coble and J. Bright (2019) Probabilistic genotyping software: an overview. Forensic Science International: Genetics 38, p. 219â224. Cited by: §2.2. M. D. Coble, J. Buckleton, J. M. Butler, T. Egeland, R. Fimmers, P. Gill, L. GusmĂŁo, B. Guttman, M. Krawczak, N. Morling, et al. (2016) DNA commission of the international society for forensic genetics: recommendations on the validation of software programs performing biostatistical calculations for forensic genetics applications. Forensic Science International: Genetics 25, p. 191â197. Cited by: §2.3. Commonwealth v. Foley, 38 A.3d 882, 888â90 ((Pa. Super. Ct. 2012)) Cited by: §2.1. S. Costanza-Chock, I. D. Raji, and J. Buolamwini (2022) Who audits the auditors? recommendations from a field scan of the algorithmic auditing ecosystem. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, p. 1571â1583. Cited by: §1, §3.2.1, §3.2, §6.2.1. N. R. Council, D. on Engineering, P. Sciences, C. on Applied, T. Statistics, G. Affairs, C. on Science, Law, and C. on Identifying the Needs of the Forensic Sciences Community (2009) Strengthening forensic science in the united states: a path forward. National Academies Press. Cited by: §A.1, §2.3. W. H. Deng, M. Nagireddy, M. S. A. Lee, J. Singh, Z. S. Wu, K. Holstein, and H. Zhu (2022) Exploring how machine learning practitioners (try to) use fairness toolkits. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, p. 473â484. Cited by: §1, §3.2. [21] (2021) Factsheet for ANSI/ASB Standard 018. Factsheet American Academy of Forensic Sciences. External Links: Link Cited by: §5.4.3, §5. T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. D. Iii, and K. Crawford (2021) Datasheets for datasets. Communications of the ACM 64 (12), p. 86â92. Cited by: §3.2. S. R. Geiger, U. Tandon, A. Gakhokidze, L. Song, and L. Irani (2024) Making algorithms public: reimagining auditing from matters of fact to matters of concern. International Journal of Communication 18, p. 634â655. External Links: Link Cited by: §6.2.3. M. K. Gerchick, R. EncarnaciĂłn, C. Tanigawa-Lau, L. Armstrong, A. GutiĂŠrrez, and D. Metaxa (2025) Auditing the audits: lessons for algorithmic accountability from local law 144âs bias audits. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, p. 29â44. Cited by: §1, §3.1, §3.2.1. E. P. Goodman and J. Trehu (2022) Algorithmic auditing: chasing ai accountability. Santa Clara High Tech. LJ 39, p. 289. Cited by: §1, §3.1, §3.1, §3.2.1, §3.2. L. Groves, J. Metcalf, A. Kennedy, B. Vecchione, and A. Strait (2024) Auditing work: exploring the new york city algorithmic bias audit regime. In The 2024 ACM Conference on Fairness, Accountability, and Transparency, p. 1107â1120. Cited by: §1, §3.2.1, §6.2.2, §6.2.2. [27] (2023-10) Guideline for Internal Validation / Verification of Various Aspects of the DNA Profiling Process. Guidelines The European Network of Forensic Science Institutes (ENFSI). External Links: Link Cited by: §2.3. [28] (2015-06) Guidelines for the Validation of Probabilistic Genotyping Systems. Guidelines Scientific Working Group on DNA Analysis Methods (SWGDAM). External Links: Link Cited by: §2.3. K. Holstein, J. Wortman Vaughan, H. DaumĂŠ I, M. Dudik, and H. Wallach (2019) Improving fairness in machine learning systems: what do industry practitioners need?. In Proceedings of the 2019 CHI conference on human factors in computing systems, p. 1â16. Cited by: §1, §1, §6.1. C. J. Hoofnagle (2016) Assessing the federal trade commissionâs privacy assessments. IEEE Security & Privacy 14 (2), p. 58â64. Cited by: §3.2.1. A. S. Hub (2024) External Links: Link Cited by: §1, §6.2.3. [32] (2018) Internal Validation of STRmix V2.4 for Fusion NYC OCME. Internal Validation Study Summary NYC Office of the Medical Examiner (OCME). External Links: Link Cited by: §4.1.2, §4.2, §5.1.1, §5.1.2, §5.2.1, §5.2.3, §5.3.1, §5.4.1, §5.5.1. [33] (2018) Internal Validation of STRmix V2.5 for the Colorado Bureau of Investigation (CBI) Forensic Laboratories (GlobalFiler, 3500xL CE). Internal Validation Study Summary Colorado Bureau of Investigation (OCME). External Links: Link Cited by: §4.1.2, §4.2, §5.1.1, §5.1.2, §5.2.3, §5.4.2. [34] (2021) Internal Validation of STRmix v2.7 for Fusion 5C/3500xL Data. Internal Validation Study Summary NYC Office of the Medical Examiner (OCME). External Links: Link Cited by: §4.1.2, §4.2, §5.1.1, §5.1.1, §5.1.2, §5.2.3, §5.3.2, §5.4.1, §5.4.2, §5.4.2, §5.5.1. S. J. Jackson and S. Barbrow (2015) Standards and/as innovation: protocols, creativity, and interactive systems development in ecology. In Proceedings of the 33rd annual ACM conference on human factors in computing systems, p. 1769â1778. Cited by: §6.2.1. A. Jin and N. Salehi (2024) (Beyond) reasonable doubt: challenges that public defenders face in scrutinizing ai in court. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, p. 1â19. Cited by: §1, §2.3.2, §4.4.2, §6.2.3, §6.2.3. H. Kelly, M. Coble, M. Kruijver, R. Wivell, and J. Bright (2022) Exploring likelihood ratios assigned for siblings of the true mixture contributor as an alternate contributor. Journal of forensic sciences 67 (3), p. 1167â1175. Cited by: §5.1.2. D. E. Krane and M. K. Philpott (2022) Using laboratory validation to identify and establish limits to the reliability of probabilistic genotyping systems. In Handbook of DNA Profiling, p. 297â319. Cited by: §4.4.2, §6.2.3. K. Lam, B. Lange, B. Blili-Hamelin, J. Davidovic, S. Brown, and A. Hasan (2024) A framework for assurance audits of algorithmic systems. In The 2024 ACM Conference on Fairness, Accountability, and Transparency, p. 1078â1092. Cited by: §1, §1, §3.2.1, §3.2, §6.2.1. M. S. Lam, A. Pandit, C. H. Kalicki, R. Gupta, P. Sahoo, and D. Metaxa (2023) Sociotechnical audits: broadening the algorithm auditing lens to investigate targeted advertising. Proceedings of the ACM on Human-Computer Interaction 7 (CSCW2), p. 1â37. Cited by: §3.1. M. Latonero and A. Agarwal (2021) Human rights impact assessments for ai: learning from facebookâs failure in myanmar. Carr Center for Human Rights Policy Harvard Kennedy School, Harvard University. Cited by: §3.1. C. Lawrence, I. Cui, and D. Ho (2023) The bureaucratic challenge to ai governance: an empirical assessment of implementation at us federal agencies. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, p. 606â652. Cited by: §4.2.1. M. S. A. Lee and J. Singh (2021) The landscape and gaps in open source fairness toolkits. In Proceedings of the 2021 CHI conference on human factors in computing systems, p. 1â13. Cited by: §1, §3.2. M. A. Madaio, L. Stark, J. Wortman Vaughan, and H. Wallach (2020) Co-designing checklists to understand organizational challenges and opportunities around fairness in ai. In Proceedings of the 2020 CHI conference on human factors in computing systems, p. 1â14. Cited by: §1, §1, §3.2, §3.2, §6.1. [45] (2024) Maryland State Police Forensic Sciences Division Internal Validation of STRmix V2.9.1). Internal Validation Study Summary Maryland State Police (MSP). External Links: Link Cited by: §4.1.2, §4.2, §5.1.2, §5.1.3, §5.2.1, §5.2.2, §5.2.3, §5.3.2. J. Matthews, M. Babaeianjelodar, S. Lorenz, A. Matthews, M. Njie, N. Adams, D. Krane, J. Goldthwaite, and C. Hughes (2019) The right to confront your accusers: opening the black box of forensic dna software. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, p. 321â327. Cited by: §4.4.2. J. N. Matthews, G. Northup, I. Grasso, S. Lorenz, M. Babaeianjelodar, H. Bashaw, S. Mondal, A. Matthews, M. Njie, and J. Goldthwaite (2020) When trusted black boxes donât agree: incentivizing iterative improvement and accountability in critical software systems. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, p. 102â108. Cited by: §4.4.2. J. Metcalf, E. Moss, E. A. Watkins, R. Singh, and M. C. Elish (2021) Algorithmic impact assessments and accountability: the co-construction of impacts. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, p. 735â746. Cited by: §3.2. M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru (2019) Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency, p. 220â229. Cited by: §1, §3.2. G. S. Morrison, C. Neumann, and P. H. Geoghegan (2020) Vacuous standardsâsubversion of the osac standards-development process. Vol. 2, Elsevier. Cited by: §2.3.2, §2.3, §3.2.1, §6.1. K. L. Moss (2015) The admissibility of trueallele: a computerized dna interpretation system. Wash. & Lee L. Rev. 72, p. 1033. Cited by: §2.1, footnote 3. D. K. Mulligan, C. Koopman, and N. Doty (2016) Privacy is an essentially contested concept: a multi-dimensional analytic for mapping privacy. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 374 (2083), p. 20160118. Cited by: §6.2.3. D. K. Mulligan, J. A. Kroll, N. Kohli, and R. Y. Wong (2019) This thing called fairness: disciplinary confusion realizing a value in technology. Proceedings of the ACM on Human-Computer Interaction 3 (CSCW), p. 1â36. Cited by: §6.2.3. Z. Obermeyer, B. Powers, C. Vogeli, and S. Mullainathan (2019) Dissecting racial bias in an algorithm used to manage the health of populations. Science 366 (6464), p. 447â453. Cited by: §1. [55] A. A. of Forensic Sciences (AAFS)About asb(Website) External Links: Link Cited by: §2.3.1. V. Ojewale, R. Steed, B. Vecchione, A. Birhane, and I. D. Raji (2025) Towards ai accountability infrastructure: gaps and opportunities in ai audit tooling. CHI â25, New York, NY, USA. External Links: ISBN 9798400713941, Link, Document Cited by: §1, §1, §1, §1, §3.1, §3.2.1, §3.2.1, §3.2, Figure 2, §4.3, 2nd item, §6.1, §6.2.1, §6.2.1, §6.2.3. [57] (2019) Palm Beach County Sheriffâs Office Internal Validation of STRmix V2.6.2 (Powerplex Fusion6C, 3500xICE). Internal Validation Study Summary Palm Beach County Sheriffâs Office (PBSO). External Links: Link Cited by: §4.1.2, §4.2, §5.1.2, §5.1.3, §5.2.3, §5.5.1. D. R. Paoletti, T. E. Doom, C. M. Krane, M. L. Raymer, and D. E. Krane (2005) Empirical analysis of the str profiles resulting from conceptual mixtures. Journal of forensic sciences 50 (6), p. JFS2004475â6. Cited by: §2.1. [59] (2025-07) Quality Assurance Standards for Forensic DNA Testing Laboratories. Guidelines Federal Bureau of Investigation (FBI). External Links: Link Cited by: §2.3. E. Radiya-Dixit and G. Neff (2023) A sociotechnical audit: assessing police use of facial rrecognition. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, p. 1334â1346. Cited by: §3.1, §6.2.1. I. D. Raji, A. Smart, R. N. White, M. Mitchell, T. Gebru, B. Hutchinson, J. Smith-Loud, D. Theron, and P. Barnes (2020) Closing the ai accountability gap: defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the 2020 conference on fairness, accountability, and transparency, p. 33â44. Cited by: §1, §1, §3.2, 2nd item, §6.2.1, §6.2.3. I. D. Raji, P. Xu, C. Honigsberg, and D. Ho (2022) Outsider oversight: designing a third party audit ecosystem for ai governance. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, p. 557â571. Cited by: §1, §1, §3.1, §3.1, §3.2.1, §3.2.1, §3.2, §6.2.1, §6.2.3. B. Rakova, J. Yang, H. Cramer, and R. Chowdhury (2021) Where responsible ai meets reality: practitioner perspectives on enablers for shifting organizational practices. Proceedings of the ACM on Human-Computer Interaction 5 (CSCW1), p. 1â23. Cited by: §1, §3.2, §6.1. A. Roth (2016) Machine testimony. Yale LJ 126, p. 1972. Cited by: §2.2.1. B. G. Schor, C. Norval, E. Charlesworth, and J. Singh (2024) Mind the gap: designers and standards on algorithmic system transparency for users. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, p. 1â16. Cited by: §1, §3.2.1, §6.1, §6.2.3. A. D. Selbst (2021) An institutional view of algorithmic impact assessments. Harv. JL & Tech. 35, p. 117. Cited by: §3.1. [67] B. D. ServicesThe kinship problem(Website) External Links: Link Cited by: §4.4.2. M. Sinha (2022) Radically reimagining forensic evidence. Ala. L. Rev. 73, p. 879. Cited by: §2.3. [69] (2024-07) Software Validation for DNA Mixture Interpretation. Guidelines UK Forensic Science Regulator. External Links: Link Cited by: §2.3. STRmix (Mar 11, 2025) External Links: Link Cited by: §2.1, §4.4.2, footnote 12. United States v. Anderson, 673 F. Supp. 3d 671 (21-CR-204) ((M.D. Pa. 2023)) Cited by: §1, footnote 10. United States v. Johnston, (23-CR-13) ((E.D. N.Y. 2025)) Cited by: §2.2.1. United States v. Lewis, 442 F. Supp. 3d 1122 (18-CR-194) ((D. Minn. 2020)) Cited by: §2.2.1. United States v. Ortiz, 736 F.Supp.3d 895 (21-CR-2503) ((S.D. Cal. 2024)) Cited by: §2.2.1. R. Y. Wong, M. A. Madaio, and N. Merrill (2023) Seeing like a toolkit: how toolkits envision the work of ai ethics. Proceedings of the ACM on Human-Computer Interaction 7 (CSCW1), p. 1â27. Cited by: §1, §3.2. L. Wright, R. M. Muenster, B. Vecchione, T. Qu, P. Cai, A. Smith, C. 2. S. Investigators, J. Metcalf, J. N. Matias, et al. (2024) Null compliance: nyc local law 144 and the challenges of algorithm accountability. In The 2024 ACM Conference on Fairness, Accountability, and Transparency, p. 1701â1713. Cited by: §3.1, §3.2.1, §3.2.1, §6.2.2, §6.2.2. E. Wu, K. Wu, R. Daneshjou, D. Ouyang, D. E. Ho, and J. Zou (2021) How medical ai devices are evaluated: limitations and recommendations from an analysis of fda approvals. Nature Medicine 27 (4), p. 582â584. Cited by: §1. Sen. R. Wyden (2023) Algorithmic accountability act of 2023, s. 2892, 118th congress. Note: https://w.congress.gov/bill/118th-congress/senate-bill/2892 Cited by: §3.1. Appendix A Additional Background A.1. The creation of the ASB The creation of the AAFS Standards Board (ASB) was in large part motivated by the 2009 National Academies report, which recommended, amongst other things, the creation of standards for forensic science to âensure desirable characteristics of services and techniques such as quality, reliability, efficiency, and consistency among practitionersâ and âmaintain autonomy from vested interest groupsâ (Council et al., 2009). As opposed to other organizations that responded to this call, the ASB was specifically created in response to interest in the forensic science community, in developing an ANSI-accredited standards-developing organization (Ambrosius, Nov 4, 2025). As the organization describes on their website, âIn July 2015, AAFS received a $1.5 million grant from the Laura and John Arnold Foundation over a four-year period to become an accredited [standards-developing organization] and provide the standards to the public free of charge. Since then, the [ASB] has formed 15 discipline-specific Consensus Bodies covering everything from DNA and toxicology to forensic nursing and wildlife forensicsâ (Ambrosius, Nov 4, 2025). Appendix B Additional Methods B.1. Our interpretations of ASB 018âs line-level requirements that relate to internal validation Below, for each ASB 018 line-level requirement that we deem relevant for internal validation studies (quoted in bold text), we provide our interpretation of the requirement. We understood two of these requirements to be tautological, or always be met, regardless of the internal validation studyâs specific practices (R4.1.3-d, R4.1.3-e). For several others, we interpreted the requirement to be met simply by the existence of the audit and audit report. For both of these groups of requirements, we explain our interpretation under the quoted requirement. For all other requirements included in our study, we phrased our interpretation as a question that we used to assess each audit report for compliance with the requirement, which we document below. In Appendix B.2, we document sections of each audit report that we take to fulfill each of our interpretations of this third group of line-level requirements. In constructing our interpretations, we drew on ASB 018âs own definitions of terms when available. Throughout, we aimed to reduce ambiguity while retaining as much of ASB 018âs language as possible. Notation: When a single numbered ASB 018 requirement contains multiple line-level requirements, we add letters after the ASB 018 requirement number to label each line-level requirement. For example, Requirement 4.1.3 contains five individual âshallâ statements, which we distinguish using letters (e.g., R4.1.3a-e). Several of the line-level requirements we include in our analysis are pulled from Annex A of ASB 018, which we denote using âannexâ (e.g., R4.1.3-annex). ⢠(R4.1) âThe laboratory shall validate a probabilistic genotyping system prior to its use for casework samples in the laboratory.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 3) â Does the report describe the laboratory as having at least some part in conducting or reviewing the internal validation study? ⢠(R4.1.1-a) âValidations shall include both developmental and internal studies.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 3) â Line-level requirement simply requires the existence of an internal validation study. ⢠(R4.1.1-b) âDevelopmental validation shall not replace internal validation.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 3) â Line-level requirement simply requires the existence of an internal validation study. ⢠(R4.1.3-a) âInternal validation studies shall address the following: accuracy, sensitivity, specificity, and precision.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 3) â (i) Does any part of the report address the softwareâs accuracy (i.e., the softwareâs ability to produce the expected likelihood ratio as calculated manually or with an alternate software program or application)? â (i) Does any part of the report address the softwareâs sensitivity (i.e., the softwareâs ability to produce inclusionary likelihood ratios for known contributors)? â (i) Does any part of the report address the softwareâs specificity (i.e., the softwareâs ability to produce exclusionary likelihood ratios for known non-contributors)? â (iv) Does any part of the report address the softwareâs precision (i.e., the variation in likelihood ratios calculated from repeated software analyses of the same input data using the same set of conditions/parameters)? ⢠(R4.1.3-b) âThese studies shall include internally generated case-type profiles of known composition that represent (in terms of number of contributors, mixture ratios, and total DNA template quantities) the range of actual casework samples intended for analysis with the system at the laboratory.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 3) â Does the report include internally generated samples of known composition that span a range of actual casework samples (in terms of the number of contributors, mixture ratios, and total DNA template quantities)? ⢠(R4.1.3-c) âStudies shall not be limited to pristine DNA samples but shall also include compromised DNA samples (e.g., low template, degraded, and inhibited samples).â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 3) â Does the report include any DNA samples that the lab either describes as âcompromisedâ or âchallengingâ, or is low template, degraded, or inhibited? ⢠(R4.1.3-d) âThe internal validation shall not exceed the scope of conditions tested in the developmental validation.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 3) â Because ASB 018âs broad definition of developmental validation could also encompass any internal validation, including the internal validation at hand, this requirement is always met. ⢠(R4.1.3-e) âCase type profiles that fall outside the range of conditions explored in the developmental validation shall require additional developmental validation studies.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 3) â Condition triggering this requirement is never met given ASB 018âs broad definition of developmental validation (see notes for R4.1.3-d). ⢠(R4.1.3-annex) âThe laboratory shall perform sufficient studies to address the variability inherent to the various aspects of DNA testing, data generation, analysis and interpretation of data and user input parameters.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 5) â Does any part of the report test the software using a set of at least two samples that captures at least one source of variability in DNA profile generation, DNA profile interpretation by users, or analysis? ⢠(R4.1.4-a) âInternal validation studies shall include evaluating user input parameters that vary run to run.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 3) â Does the report include any experiment that changes a user input? ⢠(R4.1.4-b) âThe effects of artifacts (e.g., stutter) and parameters that relate to the statistical algorithm (e.g., run time parameters for the software system that can vary from system to system) shall also be evaluated. [âŚ] [T]he specific parameters to be tested shall be determined by the laboratory.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 3) â (i) Does any part of the report assess the effect of an artifact on some aspect of system behavior? â (i) Does any part of the report assess the effect of a parameter that relates to the statistical algorithm on some aspect of system behavior? Are these parameters determined by the laboratory? ⢠(R4.1.5-a) âInternal validation studies shall also include the evaluation of multiple propositions for case type samples to aid in the development of propositions.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 3) â Does any part of the report use multiple propositions? ⢠(R4.1.5-b) âSuch studies shall also consider the effect of overestimating and underestimating the number of contributors.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 3) â Does any part of the report consider the effect of overestimating and underestimating the number of contributors on some aspect of system behavior? ⢠(R4.1.6) âFor internal validation, the laboratory shall evaluate both the appropriate sample types (i.e., number of contributors, mixture ratios, and template quantities) and the number of samples within each type to demonstrate the potential limitations and reliability of the software. The laboratory shall base this evaluation on the intended application of the software.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 3) â Does any part of the report claim that the studyâs evaluated samples are appropriate for the intended application of the software? ⢠(R4.4-annex) âAdditional validation or a performance check shall be based on the list of documented changes provided by the developer that accompany each updated version of the software installed in the laboratory.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 5) â If the report indicates that the lab has adopted a previous version of the software, does any part of the report suggest the study is based on information about documented changes provided by the developer? ⢠(R4.5) âAll validation and performance check studies conducted by the laboratory shall be documented and retained by the laboratory.â (ANSI/ASB Standard 018, 1st Ed., 2020, p. 4) â Because all of the audits we analyze in this study have been documented in the audit reports we are analyzing, this requirement is met for all audits in our study. B.2. Compliance assessment In Table 1, we document sections of each audit report that we take to fulfill each of our interpretations of ASB 018âs line-level requirements (see Appendix B.1 for our interpretations). Additional note on our assessment of PBSO internal validation study of STRmix v2.6.2 against R4.4-annex: While not explicitly describing differences between v2.6.2 and the previously validated v2.4, the PBSO v2.6.2 report does describe developer-provided documentation accompanying v2.6.2 as a guide for designing the PBSO validation of v2.6.2 and deliberately avoids re-validating certain software behaviors it deemed irrelevant to the study. Line-level Req. OCME v2.4 (2016) CBI v2.5 (2018) PBSO v2.6.2 (2019) OCME v2.7 (2021) MSP v2.9.1 (2024) 4.1 Introduction (p. 2) Introduction (p. 2) Introduction (p. 2) Introduction (p. 1) Introduction (p. 3) 4.1.3-a (i) Exp. 2 (p. 3-5) Sec. A (p. 4-5) Sec. A (p. 5-7) Exp. 2 (p. 4-6) Sec. A (p. 5-7) 4.1.3-a (i) Exp. 4 (p. 7-17) Sec. D (p. 7-32) Sec. D (p. 8-15) Exp. 4 (p. 14-25) Sec. D (p. 13-29) 4.1.3-a (i) Exp. 4 (p. 7-17) Sec. D (p. 7-32) Sec. D (p. 8-15) Exp. 4 (p. 14-25) Sec. D (p. 13-29) 4.1.3-a (iv) Exp. 6 (p. 19-22) Sec. M (p. 52-55) Sec. M (p. 25-31) Exp. 6 (p. 28-30) Sec. M (p. 64-66) 4.1.3-b (p. 2, 40) (p. 8, 56) (p. 2, 31) (p. 2, 64) (p. 3, 14) 4.1.3-c Exp. 14 (p. 38-40) Sec. L (p. 47-52) Sec. L (p. 25) Exp. 13 (p. 54-63) Sec. J (p. 48-58) 4.1.3-annex Exp. 9 (p. 24-27) Sec. M (p. 52-53) Sec. F (p. 18-22) Exp. 6 (p. 28-30) Sec. M (p. 64-66) 4.1.4-a Exp. 10 (p. 27-33) Sec. F (p. 35-38) Sec. F (p. 18-22) Exp. 9 (p. 36-44) Sec. F (p. 30-40) 4.1.4-b (i) Exp. 11 (p. 33-35) Sec. G (p. 40) Sec. G (p. 22) Exp. 10 (p. 44-48) Sec. G (p. 41-44) 4.1.4-b (i) Exp. 6 (p. 19-22) Sec. M (p. 53-55) Appendix 3 (p. 36-45) Exp. 6 (p. 28-30) Sec. M (p. 64-66) 4.1.5-a Exp. 5 (p. 18-19) Sec. E (p. 32-34) Sec. E (p. 16-17) Exp. 5 (p. 25-27) Sec. E (p. 29-30) 4.1.5-b Exp. 10 (p. 27-33) Sec. F (p. 35-38) Sec. F (p. 18-22) Exp. 9 (p. 36-44) Sec. F (p. 30-40) 4.1.6 Conclusion (p. 40) Conclusion (p. 56) Conclusion (p. 31) Conclusion (p. 64) Conclusion (p. 96) 4.4-annex N/A N/A Introduction (p. 2) Introduction (p. 1-2) N/A Table 1. For each internal validation report, we document the page numbers of the internal validation report that contain text that satisfies our interpretation of each line-level requirement. When multiple sections of a report could satisfy our interpretation of a given line-level requirement, we document the page numbers for one of the sections. All page numbers refer to the numbered pages of the internal validation report. The OCME v2.4 audit report was updated in 2019 to correct transcriptional errors and improve the clarity of tables, figures and end notesâthe original audit was completed in 2016.