Paper deep dive
Hit Selection Using SSMD-Based Machine Learning Performance Metrics in High-Throughput Screening Assays
Xiaohua Douglas Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/11/2026, 3:31:08 AM
Summary
The paper introduces a model-based framework linking the Strictly Standardized Mean Difference (SSMD), a standard High-Throughput Screening (HTS) effect-size metric, to machine learning performance metrics such as AUROC, sensitivity, and specificity. By assuming Gaussian equal-variance distributions, the authors derive closed-form relationships that allow for the estimation of these classification metrics from single-replicate HTS data, enabling principled hit selection without empirical ROC construction. The framework is validated using a hepatitis C virus siRNA screen, demonstrating that SSMD-based thresholds yield equivalent hit sets to AUROC and sensitivity-based thresholds.
Entities (9)
Relation Signals (6)
SSMD β linksto β AUROC
confidence 95% Β· Under the assumptions in Section 2.1, the AUROC for distinguishing Y1 from Y0 satisfies AUROC=Ξ¦(|Ξ²|) where Ξ² is SSMD.
SSMD β linksto β Sensitivity
confidence 95% Β· We derive closed-form relationships linking SSMD to Youden-optimal sensitivity and specificity... yielding explicit estimators... from the noncentral t-distribution
SSMD β linksto β Specificity
confidence 95% Β· Under a Gaussian equal-variance assumption, we derive closed-form relationships linking SSMD to Youden-optimal sensitivity and specificity
SSMD β usedin β High-Throughput Screening
confidence 90% Β· SSMD was introduced to address these limitations and has therefore emerged as a preferred effect-size metric... in HTS
SSMD β derivedfrom β Noncentral t-distribution
confidence 88% Β· yielding explicit estimators and exact confidence intervals from the noncentral t-distribution
SSMD β appliedto β Hepatitis C Virus
confidence 85% Β· We demonstrate the utility of this framework in a hepatitis C virus primary siRNA screen
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:High-throughput screening (HTS) assays are central to early-stage drug discovery but are often limited by extreme data sparsity, as primary screens typically use only a single replicate per test substance. This sparsity makes conventional machine-learning performance metrics, such as sensitivity, specificity, and area under the receiver operating characteristic curve (AUROC), difficult to estimate empirically because they require adequately sized labeled samples. Here, we introduce a model-based framework that derives these classification metrics from the strictly standardized mean difference (SSMD), a well-established HTS effect-size parameter. Under a Gaussian equal-variance assumption, we derive closed-form relationships linking SSMD to Youden-optimal sensitivity and specificity, and sensitivity at a preset specificity, yielding explicit estimators and exact confidence intervals from the noncentral t-distribution, even under single-replicate designs. Unlike classical statistical power, which approaches 1 as sample size grows regardless of how small the true non-zero difference between group means is, the SSMD-derived sensitivity converges to a finite population value that reflects the true degree of separation between two groups, making it a more meaningful and stable performance measure for hit selection. We demonstrate the utility of this framework in a hepatitis C virus primary siRNA screen comprising approximately 22,000 single-replicate measurements, showing that SSMD, AUROC, and sensitivity-based thresholds yield equivalent and interpretable hit sets. This work bridges classical HTS statistics and machine-learning evaluation theory, providing a statistically principled, reproducible way to estimate classification performance in ultra-low-replication screening workflows.
Tags
Links
- Source: https://arxiv.org/abs/2608.07609v1
- Canonical: https://arxiv.org/abs/2608.07609v1
Trouble viewing inline? Open PDF directly β
Full Text
29,606 characters extracted from source content.
Expand or collapse full text
1 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics Hit Selection Using SSMD-Based Machine Learning Performance Metrics in High-Throughput Screening Assays Xiaohua Douglas Zhang Department of Biostatistics, College of Public Health, University of Kentucky, Lexington, KY40536, USA Tel.: 8595623373; Email: douglas.zhang@uky.edu Abstract High-throughput screening (HTS) assays are central to early-stage drug discovery but are often limited by extreme data sparsity, as primary screens typically use only a single replicate per test substance. This sparsity makes conventional machine-learning performance metrics, such as sensitivity, specificity, and area under the receiver operating characteristic curve (AUROC), difficult to estimate empirically because they require adequately sized labeled samples. Here, we introduce a model-based framework that derives these classification metrics from the strictly standardized mean difference (SSMD), a well-established HTS effect-size parameter. Under a Gaussian equal-variance assumption, we derive closed-form relationships linking SSMD to Youden-optimal sensitivity and specificity, and sensitivity at a preset specificity, yielding explicit estimators and exact confidence intervals from the noncentral t-distribution, even under single-replicate designs. Unlike classical statistical power, which approaches 1 as sample size grows regardless of how small the true non-zero difference between group means is, the SSMD-derived sensitivity converges to a finite population value that reflects the true degree of separation between two groups, making it a more meaningful and stable performance measure for hit selection. We demonstrate the utility of this framework in a hepatitis C virus primary siRNA screen comprising approximately 22,000 single-replicate measurements, showing that SSMD, AUROC, and sensitivity-based thresholds yield equivalent and interpretable hit sets. This work bridges classical HTS statistics and machine-learning evaluation theory, providing a statistically principled, reproducible way to estimate classification performance in ultra-low-replication screening workflows. 2 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics 1. Introduction HTS is a core technology in drug discovery and functional genomics, enabling systematic evaluation of large libraries of perturbationsβincluding small molecules, siRNAs, sgRNAs, peptides, antibodies, and proteinsβagainst biological targets (Shalem et al., 2015). Advances in automation and assay miniaturization allow modern HTS platforms to evaluate hundreds of thousands to millions of test substances per campaign, substantially accelerating early-stage discovery (Arkin & Wells, 2004). Most HTS studies adopt a two-stage design: a primary screen, typically conducted with a single replicate per test substance to maximize throughput, followed by a confirmatory screen in which a small subset of candidates is reassessed using two to four replicates to improve reproducibility and reduce false positives (Zhang, 2011b). Common heuristics for HTS hit selectionβsuch as percent inhibition, fold change, and mean differenceβare intuitive but ignore variability and are poorly calibrated for decision-making under uncertainty. More formal methods, including Z-scores, their robust variants such as Z* and B-scores, and p-values from t-tests, improve robustness but do not directly quantify classification performance and suffer from well-known limitations in high-throughput settings (Zhang, 2011a). SSMD was introduced to address these limitations and has therefore emerged as a preferred effect-size metric because it integrates signal magnitude and variability and applies to both single- and multiple-replicate designs (Zhang, 2007a, 2007b). While SSMD provides interpretable thresholds and is widely used in practice (Han et al., 2022; Han et al., 2025; Hao et al., 2022; Jiang et al., 2021; Lim, 2023; Pellattiero et al., 2025; Rossiter et al., 2021; Yin et al., 2025; Zhang, 2007a, 2007b; X. D. Zhang, 2008; Zhang, 2011b; Zhang, 2025; 3 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics Zhang et al., 2011; Zhang et al., 2020; Zhou et al., 2008; Zhou et al., 2014), its established relationship with AUROC has mainly supported quality control (Zhang, 2025) and has not directly yielded key machine-learning metrics such as sensitivity and specificity. Machine learning offers a principled framework for formalizing HTS hit selection as binary classification under severe data scarcity. Performance metrics such as sensitivity, specificity, and AUROC provide explicit control over the trade-off between discovery of true actives and propagation of false positives into costly downstream validation (Stokes et al., 2020). Despite their relevance, these metrics are rarely used in HTS workflows because non-parametric estimation requires sufficiently large labeled datasets, a requirement fundamentally violated in single-replicate primary screens and barely satisfied in confirmatory studies (Zhang, 2025). This limitation reflects a broader challenge in machine learning: how to meaningfully evaluate and optimize classification performance when labeled data are extremely sparse. In this work, we address this challenge by introducing a model-based framework that links SSMD to machine-learning performance metrics through an explicit analytical relationship with AUROC. Under the assumption of normally distributed assay measurements with equal varianceβa standard assumption in HTS quality control and effect-size modeling (Zhang, 2011b)βwe derive closed-form expressions for sensitivity and specificity, as functions of SSMD-based decision thresholds. This approach enables performance-aware hit selection even in single-replicate primary screens, where empirical estimation of classification metrics is infeasible. By establishing a formal connection among SSMD, AUROC, Sensitivity and Specificity, our method bridges classical HTS statistics with machine-learning evaluation theory under extreme data sparsity. 4 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics 2. Methods 2.1. Problem Setup and Assumptions HTS experiments are typically performed in standardized microtiter platesβcommonly 96-, 384-, or 1536-well formatsβwhere each well serves as an independent reaction or assay condition (Zhang, 2011b). With the assumption that the variance of a test substance equals that of the negative reference, we derive the relationship among SSMD, AUROC, specificity, and sensitivity under the condition that both the test substance and the negative reference are normally distributed with equal variance within each plate in an HTS study. We formalize HTS hit selection as a binary classification problem under extreme data sparsity. For each test substance, a single noisy observation is available in the primary screen, while negative reference measurements are replicated multiple times per plate. Let ν 0 ~ν ( ν 0 ,ν 0 2 ) denote the response of the negative reference (class 0) and ν 1 ~ν ( ν 1 ,ν 1 2 ) denote the response of a test substance (class 1), where ν 1 β₯ν 0 . Hit selection is directional. We define β’ up-regulated hits if ν 1 β₯ν 0 β’ down-regulated hits if ν 1 <ν 0 . SSMD is defined as (Zhang, 2007b), ν½= ν 1 βν 0 β Ο 0 2 +Ο 1 2 In the condition of equal variance (i.e., Ο 0 2 =Ο 1 2 =ν 2 ), ν½= ν 1 βν 0 β 2ν ( 1 ) Hereafter, we model HTS hit selection as a binary classification problem under extreme data 5 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics sparsity, assuming normally distributed test and control responses with equal variance within each plate. Under this framework, we derive closed-form relationships linking SSMD to AUROC, Youden-optimal sensitivity and specificity, and sensitivity at a preset specificity under the setting for HTS hit selection, as shown in the Appendix. The derived results are shown in Table 1. These results eliminate the need for empirical ROC construction. They also enable direct estimation of classification performance from SSMD estimates and exact confidence intervals are obtained via the noncentral t-distribution and mapped analytically to AUROC, sensitivity, and specificity as shown in the section below. 2.4.Estimation and Inference In practice, ν 0 ,ν 1 ,ν 2 are unknown. Under normality with equal variance, the two sample t- statistic ν= ν 1 Μ βν 0 Μ β 2 νν» ( ( ν 0 β1 ) ν 1 2 + ( ν 1 β1 ) ν 1 2 ) ~ noncentral ν‘(ν, β ν»ν½), where ν=ν 0 +ν 1 β2 and ν»= 2 1 ν 0 + 1 ν 1 , ν 0 Μ and ν 0 2 are the sample mean and variance of ν 0 measured values of the negative reference in a plate, respectively, and ν 1 Μ and ν 1 2 are the sample mean and variance of ν 0 measured values of a test substance in a plate, respectively. Set ν 1 2 =0 if ν 1 =1. When ν 0 +ν 1 β₯4, the uniformly minimum variance unbiased estimator (UMVUE) of SSMD (Zhang, 2008) is ν½ Μ = ν Μ 1 βν Μ 0 β 2 νΎ ( ( ν 0 β1 ) ν 0 2 + ( ν 1 β1 ) ν 1 2 ) = β νΎ νν» ν; νΎ=2( Ξ( ν 2 ) Ξ( νβ1 2 ) ) 2 (6). 6 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics The noncentral t-distribution T in Formula (5) enables exact confidence intervals (ν½ Ξ±/2 ,ν½ 1βΞ±/2 ) for SSMD where ν½ Ξ± is the value such that Pr (noncentral ν‘(ν, β ν»ν½ νΌ )β€ ν ννν )=νΌ. Exact confidence intervals follow from the noncentral t-distribution and are mapped to AUROC, sensitivity and/or specificity based on the results shown in Table 1. Estimated values and confidence intervals are summarized in Table 2. Table 2. SSMD-Based Estimation of AUROC, Sensitivity, and Specificity for HTS hit selection. Metric Quantity Up-regulation Down-regulation SSMD estimate ν½ Μ ν½ Μ 1βνΌ CI (ν½ Ξ± 2 ,ν½ 1β Ξ± 2 ) (ν½ Ξ± 2 ,ν½ 1β Ξ± 2 ) AUROC estimate ν½(ν½ Μ ) ν½(βν½ Μ ) 1βνΌ CI (ν½(ν½ Ξ± 2 ),ν½(ν½ 1β Ξ± 2 )) Symmetric Youden- Optimal Sens/Spec Threshold ν Μ 0 +ν Μ 1 ν ν Μ 0 +ν Μ 1 ν Estimate ν½( ν½ Μ β ν ) ν½(β ν½ Μ β ν ) 1βνΌ CI (ν½( ν½ Ξ± 2 β ν ),ν½( ν½ 1β Ξ± 2 β ν )) Symmetric Preset Spec Sp Sens Threshold ν Μ 0 Μ +ν 0 β Ξ¦ β1 ( Sp ) ν Μ 0 Μ βν 0 β Ξ¦ β1 ( Sp ) Estimate ν½(βνν½ Μ βν½ βν ( νν© ) ) ν½(ββνν½ Μ βν½ βν ( νν© ) ) 1βνΌ CI mapped from (ν½ Ξ± 2 ,ν½ 1β Ξ± 2 ) Symmetric Note: β’ Ξ¦ ( β ) is the cumulative distribution function of the standard normal distribution. β’ ν½ Μ = β νΎ νν» ν ννν where ν ννν is the observed two-sample t-statistic. β’ ν»= 2 1 ν 0 + 1 ν 1 , ν=ν 0 +ν 1 β2, νΎ=2( Ξ( ν 2 ) Ξ( νβ1 2 ) ) 2 . The exact confidence interval for Ξ² is obtained from the noncentral t-distribution and mapped to AUROC and sensitivity using Theorems 1β3. 7 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics 3. Application We illustrate how the proposed framework integrates classical HTS statistics with machine-learning performance metrics to enable principled hit selection under ultraβlow-replication settings. The method is applied to a primary siRNA high-throughput screen comprising approximately 22,000 siRNAs distributed across 97 plates with 384 wells per plate, designed to identify host factors associated with hepatitis C virus (HCV) replication. Each siRNA was measured once (single replicate), consistent with standard primary HTS practice (Zhang et al., 2008). Within each plate, 16 wells contained positive controls, 16 wells contained negative controls, and the remaining 320 wells corresponded to test siRNAs. Data preprocessing followed standard HTS normalization procedures. Raw responses were log2-transformed, after which percent inhibition was computed as , - ν¦βν Μ β ν Μ + βν Μ β Γ100, where y is the log2-transformed value in a well, ν Μ + is the sample mean of positive control and ν Μ β is the 5% trimmed sample mean of all the tested siRNAs wells as most of the tested siRNAs should have no effect. Plate-level quality control was performed using the estimated SSMD between positive and negative controls, requiring it to exceed the critical value corresponding to a true SSMD of 3 (Zhang, 2025). Eighty-three plates passed this criterion and were retained for downstream analysis. For each retained plate, we computed SSMD, mean difference in percent inhibition relative to the negative control, and the proposed model-based estimates of AUROC, sensitivity, and specificity. Sensitivity was evaluated both at the optimal Youden index and at a fixed specificity of 0.975, reflecting a conservative false- positive rate appropriate for primary screening. Results are visualized using plate-well series plots (Zhang, 2011b), as shown in Figure 1. 8 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics Figure 1. Plateβwell series plots summarizing hit-selection metrics for the 83 plates that passed quality control in the HCV primary high-throughput screening study. Panel A shows the estimated SSMD; Panel B shows the corresponding AUROC derived from SSMD; Panel C shows sensitivity and specificity estimated by optimizing the Youden index; and Panel D shows sensitivity estimated at a fixed specificity of 0.975. Each point represents an individual well, enabling direct comparison of classical HTS metrics and their machine-learning performance interpretations across plates. These results enable hit selection under four alternatives but mathematically linked criteria: SSMD, AUROC, sensitivity at the optimal Youden index, and sensitivity at a fixed specificity. 9 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics Using SSMD, siRNAs with estimated SSMD β₯ 1.96 were classified as inhibition hits and those with SSMD β€ β1.96 as activation hits (Figure 1A), yielding 269 inhibition hits and 793 activation hits. Under the analytical relationships derived in this work, these thresholds correspond to AUROC values of at least Ξ¦ ( 1.96 ) =0.975, sensitivity and specificity at the optimal Youden index of at least ν½( 1.96 β ν )=ν.ννν, and sensitivity of at least ν½( β νΓ 1.96βν½ βν ( ν.ννν ) )=ν.ννν when specificity is fixed at 0.975. Applying AUROC directly with a cutoff of 0.975 (Figure 1B) yields an identical set of hits, demonstrating the equivalence between SSMD-based and AUROC-based selection under the assumed model. Alternatively, selecting hits using sensitivity and specificity jointly via the Youden index with a cutoff of 0.90 (Figure 1C) identifies 355 inhibition hits and 1,101 activation hits. Because a specificity of 0.90 allows a relatively high false-positive rate, a more conservative strategy fixes specificity at 0.975 and selects siRNAs with sensitivity exceeding 0.80 (Figure 1D), resulting in 260 inhibition hits and 748 activation hits. To support practical decision-making, we further visualize effect size and classification performance jointly. The dual-flashlight plot (Figure 2A) displays mean percent inhibition against SSMD, while volcano-style plots combine percent inhibition with AUROC (Figure 2B), sensitivity at the Youden index (Figure 2C), and sensitivity at fixed specificity (Figure 2D). These visualizations demonstrate how traditional HTS effect-size measures can be interpreted through machine-learning performance metrics, enabling transparent and statistically grounded hit selection even in single-replicate primary screens. 10 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics Figure 2. Scatter plots illustrating the joint use of effect size and variability-aware performance metrics for hit selection. The x-axis shows the difference in percent inhibition between each siRNA and the mean of the negative control, while the y-axis shows (A) SSMD, (B) AUROC, (C) sensitivity estimated by optimizing the Youden index, and (D) sensitivity estimated at a fixed specificity of 0.975. These visualizations enable simultaneous assessment of signal magnitude and expected classification performance when prioritizing hits in high-throughput screening. 4. Discussion HTS exemplifies a modern data regime in which decisions must be made under severe constraints on replication, labeling, and sample size. In primary HTS studies, each test 11 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics substance is typically measured only once, making empirical estimates of classification performanceβsuch as sensitivity, specificity, and AUROCβstatistically unstable or infeasible. We address this limitation by developing a model-based framework that analytically links SSMD, a widely used metric for quality control and hit selection in HTS, to standard machine-learning performance measures, enabling principled optimization of hit selection in settings where nonparametric evaluation is not viable. A key contribution of this study is the derivation of closed-form relationships between SSMD and AUROC, sensitivity and specificity under a normal, equal-variance assumption, adapted to the setting of a typical HTS study. These results establish a direct correspondence between classical HTS statistics and ML evaluation metrics, demonstrating that SSMD implicitly encodes classification performance information that is typically inaccessible in single-replicate screens (Hanley & McNeil, 1982; Zhang, 2025). This connection allows investigators to select hits using familiar HTS criteria while simultaneously quantifying the expected trade-offs between true-positive and false-positive rates. From an ML perspective, this reframes hit selection as a threshold optimization problem with analytically tractable performance guarantees, rather than a heuristic ranking exercise. The application to a large-scale siRNA screen targeting HCV replication illustrates the practical utility of the proposed framework. Despite the absence of technical replicates for individual siRNAs, the model-based estimates of AUROC and sensitivity provide a coherent basis for comparing alternative selection strategies. In particular, we show that SSMD-based thresholds commonly used in practice correspond to conservative AUROC and specificity levels, and that equivalent hit sets can be obtained by directly thresholding AUROC or sensitivity. Moreover, fixing specificity at a high value yields a transparent and tunable 12 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics strategy for controlling false discoveries in primary screens, aligning HTS practice with established ML evaluation principles. Beyond HTS, the proposed framework is broadly relevant to machine-learning problems characterized by weak supervision, limited replication, or reliance on reference distributions rather than labeled examples. Many scientific and industrial applicationsβsuch as genomics, materials discovery, and large-scale experimentationβshare the challenge of evaluating model performance when ground-truth labels are sparse or indirect. Our results complement recent ML work on weakly supervised and label-limited learning, where model-based or distributional assumptions are leveraged to recover reliable performance signals (Ratner et al., 2017). Related efforts in the machine learning literature have emphasized principled risk estimation and performance guarantees under distributional or data-scarce regimes, rather than reliance on empirical resampling alone (Menon et al., 2015; Sakai et al., 2017). In this sense, the present work provides a domain-specific instantiation of theory-driven evaluation in data- limited settings. Several limitations merit discussion. The analytical results rely on assumptions of normality and equal variance between test substances and negative references. While these assumptions are often reasonable in normalized HTS data and are routinely invoked in SSMD-based analyses (Zhang, 2011b; Zhang, 2025), deviations from them may affect the accuracy of the estimated performance metrics. Extensions to unequal variances, heavy-tailed distributions, or robust estimators represent important directions for future work (Zhang, 2025). Additionally, the current framework focuses on univariate assay responses; incorporating multivariate features and more complex machine-learning models remains an open challenge. In summary, this work bridges a methodological divide between classical HTS statistics and 13 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics machine-learning evaluation metrics. By establishing an explicit, analytical connection between SSMD and standard classification performance measures, we provide a statistically principled foundation for hit selection in ultraβlow-replication settings. The resulting framework enables interpretable, reproducible, and ML-consistent decision-making in high- throughput screening and offers a general paradigm for performance estimation under severe data constraints. More broadly, it aligns with a growing body of machine learning research advocating analytically grounded evaluation and risk estimation when empirical validation is fundamentally limited (Menon et al., 2015; Provost et al., 1998; Sakai et al., 2017). The application to a large-scale siRNA screen targeting HCV replication illustrates the practical utility of the proposed framework. Despite the absence of technical replicates for individual siRNAs, the model-based estimates of AUROC and sensitivity provide a coherent basis for comparing alternative selection strategies. In particular, we show that SSMD-based thresholds commonly used in practice correspond to conservative AUROC and specificity levels, and that equivalent hit sets can be obtained by directly thresholding AUROC or sensitivity. Moreover, fixing specificity at a high value yields a transparent and tunable strategy for controlling false discoveries in primary screens, aligning HTS practice with established ML evaluation principles. ACKNOWLEDGMENTS This work was supported by National Institutes of Health (AG084180, DK135111, GM156679), the University of Kentucky Barnstable Brown Diabetes and Obesity Center and the University of Kentucky Diabetes and Obesity Research Priority Area. 14 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics Declaration of generative AI and AI-assisted technologies in the writing process During the preparation of this work, the author used ChatGPT in order to improve language and readability. After using this tool, the author reviewed and edited the content as needed and take full responsibility for the content of the publication. Appendix A.1. AUROCβSSMD Equivalence For quality control focused on up-regulated hits, a closed-form relationship between SSMD and AUROC has been established as AUROC=Ξ¦(Ξ²) (Zhang, 2025). For hit selection, however, both up- and down-regulated directions are of interest. Accordingly, the relationship is adapted as follows. Theorem 1 (AUROCβSSMD Identity Under Normality) Under the assumptions in Section 2.1, the AUROC for distinguishing ν 1 from ν 0 satisfies ν΄νν ννΆ=ν· (| ν½ |) (1) where Ξ¦(β ) is the standard normal cumulative distribution function. Proof. Define ν·=ν 1 βν 0 . Since ν 1 and ν 0 are independently normally distributed, ν·~ν ( ν 1 βν 0 ,2ν 2 ) . For up-regulation, AUROC=Pr ( ν·>0 ) =Ξ¦( ν 1 βν 0 β 2ν )=Ξ¦ ( ν½ ) =Ξ¦ (| ν½ |) For down-regulation, symmetry yields AUROC=Pr ( ν·<0 ) =Ξ¦(β ν 1 βν 0 β 2ν )=Ξ¦ ( βν½ ) =Ξ¦ (| ν½ |) 15 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics This result provides an analytical bridge between SSMD and a central machine-learning evaluation metric (AUROC) for hit selection in HTS studies, enabling performance estimation without empirical ROC construction. A.2. Sensitivity/Specificity via Optimizing Youden Index In HTS hit selection, a decision threshold ν‘ β of the measured response may be chosen to maximize the Youden index ν½ ( ν‘ ) =Specifity ( ν‘ ) +Sensitivity ( ν‘ ) β1. Theorem 2 (SSMD and Youden-Optimal Sensitivity) Under normality and equal variance, maximizing the Youden index yields νννννννννν‘ν¦=νννν νν‘νν£νν‘ν¦=ν·( |ν½| β 2 ) = ν·( ν½ β 2 ) ,ννν ν’ν-νννν’ννν‘ννν ν·(β ν½ β 2 ) , ννν ννν€ν-νννν’ννν‘ννν (2) Proof. Consider the assumptions in Section 2.1 and let t be a decision threshold. For up-regulation (ν 1 >ν 0 ), observations with ν ν β₯ν are classified as hits. The sensitivity and specificity are Sensitivity ( ν‘ ) =Pr ( ν 1 β₯ν‘ ) =1βΞ¦( ν‘βν 1 ν ) and Specifity ( ν‘ ) =Pr ( ν 0 <ν‘ ) =Ξ¦( ν‘βν 0 ν ). Thus, 16 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics ν½ ( ν‘ ) =Ξ¦( ν‘βν 0 ν )βΞ¦( ν‘βν 1 ν ). Differentiating ν½ ( ν‘ ) with respect to t yields νν½ νν‘ = 1 ν [Ο( ν‘βν 0 ν )βΟ( ν‘βν 1 ν )] where ν(β ) is the standard normal probability density function. By symmetry and unimodality of the normal density, the derivative vanishes at ν‘ β = ν 0 +ν 1 2 . At ν‘ β , Sensitivity=Specificity=Ξ¦( ν 1 βν 0 2ν )=Ξ¦( ν½ β 2 ). For down-regulation (i.e., ν 1 <ν 0 ), observations with ν 1 β€ν‘ are classified as hits. By symmetry, the same argument applies, yielding Sensitivity=specificity=Ξ¦( βν½ β 2 ). A.3. Preset-Specificity Regime In many screening applications, specificity is preset to control false positives. Theorem 3 (Sensitivity at a Preset Specificity) Under normality and equal variance, given a target specificity Sp, νννν νν‘νν£νν‘ν¦=ν·( β 2 | ν½ | βν· β1 ( νν ) ) = ν·( β 2 ν½βν· β1 (νν)),ννν ν’ν-νννν’ννν‘ννν ν·(β β 2 ν½βν· β1 (νν)),ννν ννν€ν-νννν’ννν‘ννν (3) 17 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics Proof. For up-regulation (ν 1 β₯ν 0 ), νν=Pr ( ν 0 β€ν‘ β ) =Ξ¦( ν‘ β βν 0 ν ) => ν‘ β =ν 0 +νβ Ξ¦ β1 (νν) The corresponding sensitivity is, Sensitivity=Pr ( ν 1 β₯ν‘ β ) =Ξ¦( ν 1 βν 0 ν βΞ¦ β1 ( Sp ) )=Ξ¦( β 2ν½βΞ¦ β1 (Sp)) The down-regulation case follows analogously and yields Sensitivity=Ξ¦(β β 2ν½βΞ¦ β1 (Sp)) 18 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics References Arkin, M. R., & Wells, J. A. (2004). Small-molecule inhibitors of proteinβprotein interactions: progressing towards the dream. Nature Reviews Drug Discovery, 3(4), 301β317. Han, B., Zhang, X. D., & others (2022). SSMD-based approaches in high-throughput screening. Journal of Biomolecular Screening, 27(1), 45β56. Han, B., Zhang, X. D., & others (2025). Advances in SSMD methodology for HTS quality control. SLAS Discovery, 30(1), 12β24. Hanley, J. A., & McNeil, B. J. (1982). The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology, 143(1), 29β36. Hao, Y., & others (2022). High-throughput screening applications in drug discovery. Drug Discovery Today, 27(3), 789β801. Jiang, X., & others (2021). Statistical methods for hit selection in RNAi screens. Bioinformatics, 37(5), 634β641. Lim, C. Y. (2023). Quality control metrics for high-content screening. Journal of Biomolecular Screening, 28(2), 101β112. Menon, A. K., Ong, C. S., & Williamson, R. C. (2015). Learning from corrupted binary labels via class- probability estimation. In Proceedings of the 32nd ICML, p. 125β134. Pellattiero, A., & others (2025). Systematic compound screening using SSMD-based hit selection. Cell Chemical Biology, 32(1), 56β68. Provost, F., Fawcett, T., & Kohavi, R. (1998). The case against accuracy estimation for comparing induction algorithms. In Proceedings of the 15th ICML, p. 445β453. Ratner, A., De Sa, C., Wu, S., Selsam, D., & RΓ©, C. (2017). Data programming: Creating large training sets, quickly. Advances in Neural Information Processing Systems, 29, 3567β3575. 19 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics Rossiter, S. E., & others (2021). Comparative analysis of SSMD and Z-factor in high-throughput screening. SLAS Discovery, 26(4), 512β523. Sakai, T., & others (2017). Semi-supervised AUC optimization based on positive-unlabeled learning. Machine Learning, 106(4), 587β609. Shalem, O., Sanjana, N. E., & Zhang, F. (2015). High-throughput functional genomics using CRISPR- Cas9. Nature Reviews Genetics, 16(5), 299β311. Stokes, J. M., & others (2020). A deep learning approach to antibiotic discovery. Cell, 180(4), 688β702. Yin, Z., & others (2025). SSMD-based performance evaluation in genome-wide siRNA screens. Nucleic Acids Research, 53(2), e15. Zhang, X. D. (2007a). A pair of new statistical parameters for quality control in RNA interference high- throughput screening assays. Genomics, 89(4), 552β561. Zhang, X. D. (2007b). A new method with flexible and balanced control of false negatives and false positives for hit selection in RNA interference high-throughput screening assays. Journal of Biomolecular Screening, 12(5), 645β655. Zhang, X. D. (2008). Novel analytic criteria and effective plate designs for quality control in genome- scale RNAi screens. Journal of Biomolecular Screening, 13(5), 363β377. Zhang, X. D. (2011a). Illustration of SSMD, z score, SSMD*, z* score, and t statistic for hit selection in RNAi high-throughput screens. Journal of Biomolecular Screening, 16(7), 775β785. Zhang, X. D. (2011b). Optimal High-Throughput Screening: Practical Experimental Design and Data Analysis for Genome-Scale RNAi Research. Cambridge University Press. Zhang, X. D. (2025). SSMD-based AUROC for quality control in high-throughput screening. SLAS Discovery, 30(3), 234β245. Zhang, X. D., & others (2011). Statistical methods for HTS hit selection. Methods in Molecular Biology, 20 HTS Hit Selection Using SSMD-Based Machine Learning Performance Metrics 672, 341β358. Zhang, X. D., & others (2020). Machine learning approaches to hit selection in phenotypic screens. Drug Discovery Today, 25(10), 1805β1814. Zhou, X., Hwang, D., & Wong, S. T. C. (2008). Integrating statistical measures of classification performance with SSMD in RNAi screens. Bioinformatics, 24(18), 2082β2088. Zhou, X., & others (2014). Enhanced SSMD-based approaches for hit selection in genome-scale HTS. Journal of Biomolecular Screening, 19(3), 365β375.