Paper deep dive
Trustworthy Protein-Ligand Binding Affinity Prediction via Reliability-Aware Multi-Engine Fusion
Yongchan Hong, Defu Cao, Wenjin Liu, Thomas Ku, Jordy Homing Lam, Emily Nguyen, Willie Neiswanger, Vsevolod Katritch, Yan Liu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/21/2026, 5:19:57 AM
Summary
The paper introduces RELIABLE-BA, a framework for protein-ligand binding affinity prediction that combines evidential uncertainty quantification with context-dependent reliability modeling. It models multiple docking engines as evidential experts using Normal-Inverse-Gamma distributions, scales their epistemic uncertainty based on learned reliability from molecular context, and fuses them via closed-form Mixture of Normal-Inverse-Gamma aggregation. This approach improves uncertainty calibration and reduces prediction error by up to 25% through selective filtering of high-confidence pairs.
Entities (11)
Relation Signals (9)
Yongchan Hong ā affiliatedwith ā University of Southern California
confidence 95% Ā· Yongchan Hong... Department of Quantitative & Computational Biology University of Southern California
RELIABLE-BA ā evaluatedon ā BDB2020+
confidence 95% Ā· Experiments on the PDBBind and BDB2020+ benchmarks demonstrate competitive point prediction
RELIABLE-BA ā evaluatedon ā PDBBind
confidence 95% Ā· Experiments on the PDBBind and BDB2020+ benchmarks demonstrate competitive point prediction
RELIABLE-BA ā uses ā Normal-Inverse-Gamma
confidence 95% Ā· modeling each engine as an evidential expert via Normal-Inverse-Gamma distributions
RELIABLE-BA ā uses ā MoNIG
confidence 95% Ā· fusing experts through closed-form aggregation... Mixture of NormalāInverse-Gamma (MoNIG) aggregation
RELIABLE-BA ā appliedto ā 5HT2A receptor
confidence 90% Ā· additional validation on the SARS-CoV-2 Mpro dataset and 5HT2A receptor demonstrates applicability
RELIABLE-BA ā appliedto ā SARS-CoV-2 Mpro
confidence 90% Ā· additional validation on the SARS-CoV-2 Mpro dataset... demonstrates applicability
RELIABLE-BA ā usesinputfrom ā ESM-2
confidence 85% Ā· shared representation h(x) is extracted by concatenating frozen ESM-2 protein embeddings
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Accurate protein-ligand binding affinity prediction is central to computational drug discovery, yet modern docking engines frequently disagree without indicating which prediction to trust. Consensus scoring and ensemble methods improve mean accuracy but treat all predictions identically without interpretable confidence measures or uncertainty decomposition, ignoring the chemical context of each protein-ligand pair. To address this limitation, we introduce RELIABLE-BA (RELIABiLity-aware Evidential fusion for Binding Affinity), an evidential framework for multi-engine binding affinity prediction. Our model comprises three steps: (1) modeling each engine as an evidential expert via Normal-Inverse-Gamma distributions, (2) scaling epistemic uncertainty through learned reliability from molecular context while preserving each expert's predictive mean, and (3) fusing experts through closed-form aggregation that captures both individual uncertainty and inter-engine disagreement. Experiments on the PDBBind and BDB2020+ benchmarks demonstrate competitive point prediction with substantially improved uncertainty calibration, and additional validation on the SARS-CoV-2 Mpro dataset and 5HT2A receptor demonstrates applicability to clinically relevant drug targets. Crucially, these uncertainty estimates enable reliable filtering of protein-ligand pairs, reducing prediction error by up to 25% when retaining only high-confidence pairs. To our knowledge, RELIABLE-BA is the first multi-engine binding affinity prediction framework to combine evidential fusion with context-dependent reliability, offering a principled path toward trustworthy AI-guided drug discovery. Our code is publicly available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2607.17601v1
- Canonical: https://arxiv.org/abs/2607.17601v1
Trouble viewing inline? Open PDF directly ā
Full Text
66,384 characters extracted from source content.
Expand or collapse full text
Trustworthy ProteināLigand Binding Affinity Prediction via Reliability-Aware Multi-Engine Fusion Yongchan Hong 0009-0009-8866-1690 Department of Quantitative & Computational BiologyUniversity of Southern CaliforniaLos AngelesCaliforniaUSA hongyong@usc.edu , Defu Cao 0000-0003-0240-3818 Department of Computer ScienceUniversity of Southern CaliforniaLos AngelesCaliforniaUSA defucao@usc.edu , Wenjin Liu 0000-0003-1909-8135 Department of Quantitative & Computational BiologyUniversity of Southern CaliforniaLos AngelesCaliforniaUSA wenjinl@usc.edu , Thomas Ku 0009-0006-1647-2979 Department of Quantitative & Computational BiologyUniversity of Southern CaliforniaLos AngelesCaliforniaUSA kut@usc.edu , Jordy Homing Lam 0000-0002-5496-6228 Department of ChemistryUniversity of California, BerkeleyBerkeleyCaliforniaUSA jhml@berkeley.edu , Emily Nguyen 0000-0003-4917-7336 Department of Computer ScienceUniversity of Southern CaliforniaLos AngelesCaliforniaUSA emilyn98@usc.edu , Willie Neiswanger 0000-0002-9619-5572 Department of Computer ScienceUniversity of Southern CaliforniaLos AngelesCaliforniaUSA neiswang@usc.edu , Vsevolod Katritch 0000-0003-3883-4505 Department of Quantitative & Computational BiologyUniversity of Southern CaliforniaLos AngelesCaliforniaUSA katritch@usc.edu and Yan Liu 0000-0002-7055-9518 Department of Computer ScienceUniversity of Southern CaliforniaLos AngelesCaliforniaUSA yanliu.cs@usc.edu Abstract. Accurate proteināligand binding affinity prediction is central to computational drug discovery, yet modern docking engines frequently disagree without indicating which prediction to trust. Consensus scoring and ensemble methods improve mean accuracy but treat all predictions identically without interpretable confidence measures or uncertainty decomposition, ignoring the chemical context of each proteināligand pair. To address this limitation, we introduce RELIABLE-BA (RELIABiLity-aware Evidential fusion for Binding Affinity), an evidential framework for multi-engine binding affinity prediction. Our model comprises three steps: (1) modeling each engine as an evidential expert via NormalāInverse-Gamma distributions, (2) scaling epistemic uncertainty through learned reliability from molecular context while preserving each expertās predictive mean, and (3) fusing experts through closed-form aggregation that captures both individual uncertainty and inter-engine disagreement. Experiments on the PDBBind and BDB2020+ benchmarks demonstrate competitive point prediction with substantially improved uncertainty calibration, and additional validation on the SARS-CoV-2 Mpro dataset and 5HT2A receptor demonstrates applicability to clinically relevant drug targets. Crucially, these uncertainty estimates enable reliable filtering of protein-ligand pairs, reducing prediction error by up to 25% when retaining only high-confidence pairs. To our knowledge, RELIABLE-BA is the first multi-engine binding affinity prediction framework to combine evidential fusion with context-dependent reliability, offering a principled path toward trustworthy AI-guided drug discovery. Our code is publicly available at https://github.com/yongchand/RELIABLE-BA. Binding Affinity Prediction, Uncertainty Quantification, Evidential Regression, Computational Drug Discovery ā ccs: Computing methodologies Artificial intelligenceā ccs: Computing methodologies Uncertainty quantificationā ccs: Applied computing Bioinformatics Figure 1. Pairwise error correlations between four binding affinity prediction engines on PDBbind. Prediction error is defined as the absolute difference between predicted and experimental binding affinity (pāKdpK_d/pāKipK_i). Spearman correlation coefficients range from 0.14 to 0.33. Error correlation matrix for binding affinity prediction engines. 1. Introduction Proteināligand binding affinity prediction is a core component of computational drug discovery, enabling efficient hit discovery and substantially reducing the number of candidate compounds for experimental evaluation (22). While physics-based methods such as free energy perturbation (FEP) achieve high accuracy, they require extensive sampling and are computationally prohibitive for large-scale screening (42). Machine learning (ML) approaches offer faster inference with competitive accuracy (7). Boltz-2 approaches FEP-level accuracy with Pearson correlations of 0.62ā0.66 on standard benchmarks (34), demonstrating that deep learning can meaningfully accelerate therapeutic discovery (33). Modern ML-based affinity prediction engines demonstrate distinct strengths. Physics-based methods capture fundamental energetics, structure-based models exploit geometric features, and sequence-based approaches generalize across protein families without requiring 3D structures (29, 21, 26). Nevertheless, proteināligand binding is inherently complex, involving subtle interplay of electrostatics, solvation, entropy, and conformational dynamics that no single model fully captures (5). As a result, even state-of-the-art engines exhibit complementary failure modes: Figure 1 shows that pairwise error correlations between engines are weak (Spearman Ļ = 0.14-0.33), indicating that different engines succeed on different subsets of complexes. This complementarity suggests that fusing existing engines offers a more efficient path to robust predictions than building a single universal model from scratch. Building on this complementarity strength, practitioners employ multiple docking tools to improve robustness (3, 10). Existing multi-engine strategies fall into two paradigms: consensus scoring, which combines predictions using heuristic rules such as rank aggregation or score normalization (46, 32), and ensemble learning, which aggregates outputs via averaging or meta-learning (2, 30, 35). While these approaches can improve mean accuracy, they are primarily designed to optimize aggregate performance, rather than to reason about engine-specific reliability for individual proteināligand complexes. To reason about which predictions to rely on in a given proteināligand complex, uncertainty quantification (UQ) characterizes predictive confidence rather than producing point estimates alone. This is essential for trustworthy decision-making in drug discovery, where models are frequently applied to novel scaffolds or targets outside the training distribution (36). Calibrated uncertainty estimates enable practitioners to prioritize high-confidence predictions for experimental validation, thereby reducing costly false positives. In active learning settings, uncertainty guides the selection of informative compounds for iterative screening (14). A key distinction is between aleatoric uncertainty, arising from inherent measurement noise that cannot be reduced, and epistemic uncertainty, reflecting limited model knowledge that could be addressed with additional data or improved models (19, 6). Without this decomposition, practitioners cannot determine whether low confidence stems from inherent target difficulty or correctable gaps in model coverage. Recent advances in uncertainty quantification provide principled frameworks for single-model settings. Evidential regression using NormalāInverse-Gamma (NIG) distributions enables closed-form uncertainty decomposition (1), and has shown promise in drug discovery applications (47, 45). However, these methods address uncertainty within a single model and do not account for the additional epistemic uncertainty arising from disagreement among engines. Binding affinity engines are not exchangeable as they rely on different assumptions, scoring functions, and training data with varying coverage (44). Existing single-model UQ methods cannot capture this inter-engine disagreement and determine which engine to trust for a given proteināligand complex. To address the above mentioned challenges, we introduce RELIABLE-BA (RELIABiLity-aware Evidential fusion for Binding Affinity). RELIABLE-BA models each docking engine as an evidential expert outputting a NIG distribution that captures both aleatoric and epistemic uncertainty. A reliability network learns context-dependent trustworthiness scores conditioned on proteināligand features, which modulate each expertās epistemic uncertainty through a principled scaling transformation while preserving predictive means. The reliability-scaled experts are then fused via closed-form Mixture of NormalāInverse-Gamma (MoNIG) aggregation (27), yielding calibrated predictions that reflect both individual engine uncertainty and inter-engine disagreement. We evaluate RELIABLE-BA on PDBbind (43) using a time-split protocol, with additional validation on BDB2020+ (24), an independent test set with zero training overlap, and a case study on the SARS-CoV-2 Mpro dataset and 5HT2A receptor. Across all settings, RELIABLE-BA achieves competitive point prediction while substantially improving uncertainty calibration. Selective prediction analyses demonstrate consistent error reduction at low coverage, and epistemic uncertainty correlates with inter-engine disagreement while point predictions remain stable. These properties make RELIABLE-BA well-suited for multi-engine virtual screening pipelines where calibrated uncertainty is needed to prioritize high-confidence predictions before committing experimental resources. In summary, this paper makes the following contributions: (1) We propose RELIABLE-BA, the first multi-engine binding affinity framework combining evidential uncertainty quantification with context-dependent reliability modeling. (2) We introduce a reliability-scaling mechanism that modulates epistemic uncertainty while preserving each expertās predictive mean, enabling coherent fusion via MoNIG aggregation. (3) Through experiments on PDBbind, BDB2020+, and both SARS-CoV-2 Mpro dataset and 5HT2A receptor case study, we demonstrate superior uncertainty calibration and up to 25.7% MAE reduction via selective prediction. Together, these results position RELIABLE-BA as a practical approach to uncertainty-aware virtual screening in drug discovery. Accurate binding affinity prediction with calibrated uncertainty is particularly critical for challenging therapeutic targets such as GPCRs, which account for over 34% of FDA-approved drug targets yet remain difficult to model due to their conformational flexibility and limited representation in binding affinity training data (15). Our 5HT2A receptor case study confirms that RELIABLE-BA maintains strong performance on this clinically relevant target class. To prove its usage in real world, we further validated the method on a SARS-CoV-2 Mpro dataset, a clinically meaningful dataset known to be independent of the training data (24). 2. Related Work 2.1. Binding Affinity Prediction Binding affinity quantifies the strength of interaction between a protein and a ligand, typically expressed as dissociation constants (KdK_d, KiK_i) or their negative logarithms (pKdK_d, pKiK_i) (23). Accurate prediction is essential for computational drug discovery, as experimental validation is costly and time-consuming, allowing only a small fraction of candidate compounds to be tested (22). Existing approaches span a wide spectrum of modeling assumptions. Physics-based approaches include molecular docking with empirical scoring functions, such as AutoDock Vina (41) and Glide (11), as well as free energy perturbation (FEP) (42). While docking methods enable efficient large-scale screening, FEP achieves substantially higher accuracy (ā¼ 1 kcal/mol error) at the cost of extensive sampling and prohibitive computational expense. Machine learning approaches offer faster inference with competitive accuracy. Structure-based methods include 3D convolutional networks that operate on voxelized proteināligand complexes (GNINA (29), KDEEP (17), Pafnucy (40)). Sequence-based models such as BIND (21) combine protein language model embeddings with ligand graph neural networks, avoiding dependence on explicit 3D structures. More recently, generative docking methods based on diffusion (DynamicBind (26)) and flow matching (FlowDock (31)) jointly predict binding poses and affinities. Despite this diversity, each approach carries distinct biases. Physics-based docking relies primarily on structural information, making it less dependent on ligand data coverage but limited in capturing complex energetics. Structure-based ML models are sensitive to pose quality, while sequence-based models may miss geometric interactions between protein and ligand (21). Critically, most methods output only point estimates without calibrated confidence scores, leaving practitioners unable to distinguish reliable predictions from unreliable ones. 2.2. Aggregation of Affinity Prediction Engines A common strategy for improving binding affinity prediction is to aggregate outputs from multiple engines. Prior work has explored two main paradigms: consensus scoring and ensemble learning, motivated by the observation that no single engine consistently outperforms others across all target classes (44). Consensus scoring aggregates predictions from independent docking or scoring functions using heuristic rules such as rank-based aggregation or score normalization (46, 32, 3). Ericksen et al. showed that machine learning consensus scoring outperforms naive combinations, but also found that the best-performing engine varies by target, making it difficult to know which engines to trust a priori. Ensemble learning formulates aggregation within a machine learning framework. Ballester and Mitchell pioneered machine learning scoring functions using Random Forests (2), while recent work has explored deep learning ensembles (30) and confidence estimation through ensemble disagreement (35). Recent advances in uncertainty quantification offer principled alternatives (36). Evidential methods have demonstrated practical impact: Soleimany et al. (38) showed that evidential uncertainty improved virtual screening hit rates, and EviDTI (47) validated uncertainty-guided predictions with experimental binding assays. UAMRL employed Normal-Inverse-Gamma distributions for uncertainty-aware fusion of multiple modalities within a single model (45). Despite these advances, existing methods either apply uniform weighting across engines or fuse modalities within a single model, without learning context-dependent reliability for each prediction source. In practice, different engines exhibit sample-level biases, which is a gap RELIABLE-BA can address through learned reliability scaling. 2.3. Uncertainty Quantification and Evidential Learning Uncertainty quantification (UQ) characterizes predictive confidence rather than producing point estimates alone. A key distinction is between aleatoric uncertainty, arising from inherent data noise, and epistemic uncertainty, reflecting limited model knowledge (19). Common approaches of UQ include deep ensembles (20), Monte Carlo Dropout (12), SWA-Gaussian (28), and sparse variational Gaussian processes (16). Another common approach is evidential regression, which is a probabilistic framework in which a neural network predicts the parameters of a NormalāInverse-Gamma (NIG) prior over the mean and variance of a Gaussian likelihood, yielding closed-form uncertainty decomposition without sampling or ensembles (1). Ma et al. (27) introduced the Mixture of NIG (MoNIG) framework for multi-source evidential fusion, combining NIG distributions through a closed-form summation operator that accumulates both within-expert uncertainty and inter-expert disagreement. While MoNIG provides a principled aggregation rule, existing applications rely on static weighting schemes that do not adapt to input-dependent predictor reliability. 3. RELIABLE-BA: RELIABiLity-aware Evidential fusion for Binding Affinity Figure 2. Overview of the RELIABLE-BA architecture. Evidential heads map each engineās score to NIG parameters, which are modulated by context-dependent reliability scores and fused via MoNIG aggregation. An example prediction illustrates per-engine reliability weights (darker = higher reliability) alongside decomposed uncertainty estimates. Architecture of RELIABLE-BA. 3.1. Problem Setup We consider the problem of predicting proteināligand binding affinity by integrating outputs from multiple binding affinity prediction engines. Let x denote a proteināligand complex, and yāāy the ground-truth binding affinity. Assume M docking engines are available, each producing a scalar score siā(x)s_i(x) for engine i. In addition, a shared representation ā(x)āādh(x) ^d is extracted by concatenating frozen ESM-2 protein embeddings (25) and ChemBERTa ligand embeddings (8). The goal is to output a predictive distribution pā(yā£x)p(y x) that captures both the predicted value and its associated uncertainty, rather than producing a single point estimate. 3.2. Evidential Expert Modeling Each engine is modeled as an evidential expert that outputs a probabilistic prediction over binding affinity. For engine i, we assume a Gaussian likelihood (1) yā£Ī¼,Ļ2ā¼ā(μ,Ļ2),y μ,Ļ^2 (μ,Ļ^2), with a NormalāInverse-Gamma (NIG) prior (1) placed over the unknown mean and variance, (2) (μ,Ļ2)ā¼NIGā(γ,ν,α,β).(μ,Ļ^2) (γ,ν,α,β). Rather than treating these as fixed parameters, we learn the NIG parameters via a neural network. Specifically, for each engine i, an MLP fĪø(i)f_Īø^(i) maps the engine score siā(x)s_i(x) to the four evidential parameters: (3) (γi,νi,αi,βi)=fĪø(i)ā(siā(x)).( _i, _i, _i, _i)=f_Īø^(i) (s_i(x) ). We enforce νi>0 _i>0 and βi>0 _i>0 via softplus activation, and αi>1 _i>1 via αi=1+softplusā(ā ) _i=1+softplus(Ā·) to ensure the expected variance ā[Ļ2]E[Ļ^2] remains finite. The four parameters admit the following interpretation: γi _i represents the predicted mean, νi _i controls confidence in the mean estimate (analogous to pseudo-observation count), while αi _i and βi _i govern the shape and scale of uncertainty over the variance. From these parameters, the expected prediction and uncertainty components are (4) ā[μ] [μ] =γi, = _i, (5) ā[Ļ2] [Ļ^2] =βiαiā1, = _i _i-1, (6) Varā(μ) (μ) =βiνiā(αiā1). = _i _i( _i-1). We interpret ā[μ]E[μ] as the point prediction, ā[Ļ2]E[Ļ^2] as aleatoric uncertainty (irreducible data noise), and Varā(μ)Var(μ) as epistemic uncertainty (model uncertainty due to limited evidence). 3.3. Reliability-Aware Uncertainty Modeling Network Since each docking engine operates on different mechanisms with unique inductive biases, context-dependent reliability will differ across proteināligand complexes. To model this behavior, RELIABLE-BA introduces a reliability network that estimates predictive trustworthiness conditioned on the input. The reliability network takes the shared embedding ā(x)h(x) as input and returns a reliability score for each engine, (7) riā(x)=Ļā(gĻā(ā(x))i),riā(x)ā(0,1),r_i(x)=Ļ (g_Ļ(h(x))_i ), r_i(x)ā(0,1), where gĻg_Ļ is an MLP and Ļ denotes the sigmoid function. The output riā(x)r_i(x) reflects how trustworthy engine i is for the given complex. The scores quantify predictive trustworthiness conditioned on molecular context, reflecting factors such as protein family, binding pocket geometry, and structural complexity. Scaling The reliability scores are used to modulate epistemic uncertainty without altering each expertās predictive bias. Specifically, RELIABLE-BA applies reliability scaling to the evidential parameters as (8) γ~i γ_i =γi, = _i, (9) ν~i ν_i =riā(x)āνi, =r_i(x)\, _i, (10) α~i α_i =riā(x)āαi+(1āriā(x)), =r_i(x)\, _i+ (1-r_i(x) ), (11) β~i β_i =riā(x)āβi. =r_i(x)\, _i. This transformation preserves each expertās predicted mean while reducing its effective evidence when reliability is low. Consequently, epistemic uncertainty increases for unreliable experts, while the expected aleatoric variance remains unchanged under the NIG parameterization. We ensure numerical validity by constraining riā(x)r_i(x) away from zero, guaranteeing ν~i>0 ν_i>0, β~i>0 β_i>0, and α~i>1 α_i>1. 3.4. MoNIG Aggregation The reliability-scaled experts are fused using a Mixture of NormalāInverse-Gamma (MoNIG) aggregation. Let (γ~i,ν~i,α~i,β~i)( γ_i, ν_i, α_i, β_i) denote the reliability-scaled evidential parameters produced by engine i. Following prior work of MoNIG (27), we aggregate multiple NormalāInverse-Gamma distributions using the NIG summation operator, denoted by ā , which defines a closed-form fusion rule that preserves conjugacy. The fusion of M scaled experts is given by (12) NIGā(γ,ν,α,β)=āØi=1MNIGā(γ~i,ν~i,α~i,β~i).NIG(γ,ν,α,β)= _i=1^MNIG( γ_i, ν_i, α_i, β_i). The fused mean is computed as an evidence-weighted average, (13) γ=āi=1Mν~iāγ~iāi=1Mν~i,γ= _i=1^M ν_i γ_i _i=1^M ν_i, with total evidence (14) ν=āi=1Mν~i.ν= _i=1^M ν_i. The remaining NormalāInverse-Gamma parameters are aggregated according to the NIG summation operator. Specifically, the fused shape parameter is (15) α=āi=1Mα~i+Mā12,α= _i=1^M α_i+ M-12, and the fused scale parameter is (16) β=āi=1Mβ~i+12āāi=1Mν~iā(γ~iāγ)2.β= _i=1^M β_i+ 12 _i=1^M ν_i ( γ_i-γ )^2. This aggregation accumulates both within-expert uncertainty and disagreement among expert means, with the second term in Eq. (16) increasing epistemic uncertainty when engines provide conflicting predictions. While the correctness of the NIG summation operator without reliability scaling has been established in prior work, we provide a formal proof in Appendix A that the summation remains valid under the proposed reliability scaling. Given the fused NormalāInverse-Gamma parameters, predictive uncertainty decomposes into aleatoric and epistemic components as (17) ā[Ļ2]āaleatoric=βαā1,Varā(μ)āepistemic=βνā(αā1). E[Ļ^2]_aleatoric= βα-1, Var(μ)_epistemic= βν(α-1). 3.5. Joint Evidential Learning Objective RELIABLE-BA is trained end-to-end with joint supervision at both the expert and fused levels (Figure 2). For each complex x with ground-truth binding affinity y, engine-specific branches output evidential parameters miā(x)=(γi,νi,αi,βi)m_i(x)=( _i, _i, _i, _i). We define the overall training objective as the sum of expert-level evidential losses and the fused-level evidential loss: (18) āā(w)=āi=1Māiā(w)+āMoNIGā(w),L(w)= _i=1^ML_i(w)+L_MoNIG(w), where each expert loss is given by (19) āiā(w)=āiNLLā(w)+Ī»āāiRā(w),L_i(w)=L^NLL_i(w)+Ī»\,L^R_i(w), and the fused loss is (20) āMoNIGā(w)=āMoNIGNLLā(w)+Ī»āāMoNIGRā(w).L_MoNIG(w)=L^NLL_MoNIG(w)+Ī»\,L^R_MoNIG(w). Here, āiNLLL^NLL_i and āMoNIGNLLL^NLL_MoNIG denote the negative log marginal likelihoods under the expert-level and fused Student-t predictive distributions, respectively. The regularization terms āiRL^R_i and āMoNIGRL^R_MoNIG follow standard evidential regression and penalize excessive evidence when prediction errors are large. The coefficient Ī» controls the trade-off between data fit and uncertainty calibration. 4. Experiments We organize our empirical analysis around the following research questions (RQs): ⢠RQ1: Does RELIABLE-BA show competitive point prediction accuracy over individual docking engines and aggregation baselines? ⢠RQ2: Does RELIABLE-BA produce better-calibrated and more informative uncertainty estimates? ⢠RQ3: Does RELIABLE-BA support risk-aware decision making through selective prediction? ⢠RQ4: How do reliability modeling and expert disagreement contribute to RELIABLE-BAās performance? 4.1. Dataset and Experimental Setup 4.1.1. Dataset and Evaluation Protocol We evaluate RELIABLE-BA on the PDBbind 2020 benchmark (43) using a time-based split to prevent temporal leakage, which was originally curated by prior works (26, 31, 39). Within this split, we exclude IC50 affinity measurements, as IC50 values are assay-dependent and not directly comparable to thermodynamic binding constants (KiK_i/KdK_d) without knowledge of assay conditions (18, 23). We have also utilized BDB2020+ (24), which is an independent, challenging dataset for fair benchmarking. Dataset details are included in Appendix B. Protein and ligand inputs are converted into fixed-dimensional representations using ESM-2 (esm2_t6_8M_UR50D) (25) for proteins and ChemBERTa (ChemBERTa-77M-MTR) (8) for ligands. All methods operate on identical 704-dimensional proteināligand embeddings. 4.1.2. Evidential Expert Models RELIABLE-BA integrates four docking engines as evidential experts, selected to maximize methodological diversity and complementary inductive biases: ⢠GNINA (Version 1.3): A Vina-derived docking engine that performs Markov chain Monte Carlo (MCMC) pose sampling followed by CNN-based rescoring on grid-based proteināligand representations (29). ⢠BIND (Version 1.4): A sequence-based affinity predictor that combines a ligand graph neural network with protein language model embeddings (ESM-2) via cross-attention (21). ⢠FlowDock (Version 0.0.3): A geometric generative docking method based on conditional flow matching, learning continuous transformations from apo to holo proteināligand complexes (31). ⢠DynamicBind (Version 1.0): An E(3)-equivariant diffusion-based docking model that learns smooth energy landscapes for ligand-specific binding (26). These experts differ substantially in input modality, architectural assumptions, and algorithms, making them well-suited for uncertainty-aware fusion. Note that the framework is modular and additional non-redundant docking engines can be incorporated as new evidential experts without architectural changes. 4.1.3. Baselines We compare RELIABLE-BA against a comprehensive set of uncertainty-aware baselines, including: (i) Deterministic MLP regression, (i) Single evidential regression (NIG) (1) (i) heteroscedastic Gaussian regression (19), (iv) MC Dropout (12), (v) Deep Ensembles with meanāvariance estimation (20), (vi) SWA-Gaussian (SWAG) (28), and (vii) Sparse Variational Gaussian Process (SVGP) (16). We have also compared with traditional aggregation methods including consensus scoring, ensemble scoring, learned weight ensemble (LWE), and softmax mixture of experts (Softmax MoE). Detailed architectural choices and hyperparameter settings for each method are provided in Appendix C. 4.1.4. Evaluation Metrics We assess performance along two complementary dimensions: Point prediction accuracy: Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), Pearson correlation, and coefficient of determination (R2R^2). Uncertainty quality: Expected Calibration Error (ECE), Continuous Ranked Probability Score (CRPS), and Negative Log-Likelihood (NLL). Uncertainty quality metrics were calculated using Uncertainty Toolbox (9). 4.2. Experimental Results Table 1. Point prediction and uncertainty evaluation on PDBbind time-split benchmark. Results averaged over 20 seeds (± standard deviation). Best results in bold, second best underlined among aggregation methods. Method MAE ā RMSE ā Corr ā R^2 ā ECE ā CRPS ā NLL ā Individual Scoring Engines DynamicBind 0.8095 1.0840 0.8507 0.6806 - - - FlowDock 1.0155 1.2776 0.7590 0.5563 - - - GNINA 1.2609 1.5621 0.6646 0.3367 - - - BIND 1.3898 1.7417 0.5415 0.1754 - - - Traditional Aggregation Methods Consensus Scoring 0.9356 1.1896 0.7912 0.6153 - - - Ensemble Scoring 0.9231 1.1533 0.8148 0.6384 - - - LWE 0.8380 1.0750 0.8350 0.6860 - - - Softmax MoE 0.8610 ± 0.0200 1.0890 ± 0.0180 0.8260 ± 0.0050 0.6780 ± 0.0110 - - - Learned Aggregation Methods Baseline 0.8630 ± 0.0287 1.0991 ± 0.0360 0.8267 ± 0.0107 0.6713 ± 0.0214 - - - Gaussian 0.8543 ± 0.0230 1.0961 ± 0.0303 0.8309 ± 0.0085 0.6732 ± 0.0182 0.0376 ± 0.0164 0.6153 ± 0.0164 1.5272 ± 0.0291 MC Dropout 0.8435 ± 0.0242 1.0796 ± 0.0256 0.8290 ± 0.0081 0.6830 ± 0.0151 0.1009 ± 0.0211 0.6286 ± 0.0145 1.5893 ± 0.0249 Deep Ensemble 0.8444 ± 0.0246 1.0887 ± 0.0324 0.8344 ± 0.0057 0.6775 ± 0.0196 0.0540 ± 0.0130 0.6134 ± 0.0165 1.5290 ± 0.0243 SVGP 0.8680 ± 0.0325 1.1162 ± 0.0320 0.8184 ± 0.0094 0.6610 ± 0.0195 0.0810 ± 0.0311 0.6416 ± 0.0166 1.5971 ± 0.0310 SWAG 0.8549 ± 0.0241 1.0961 ± 0.0281 0.8285 ± 0.0090 0.6732 ± 0.0169 0.0395 ± 0.0147 0.6155 ± 0.0151 1.5271 ± 0.0248 Evidential Regression 0.8595 ± 0.0297 1.1109 ± 0.0359 0.8279 ± 0.0062 0.6642 ± 0.0221 0.0518 ± 0.0265 0.6227 ± 0.0196 1.5381 ± 0.0335 Ours RELIABLE-BA 0.8247 ± 0.0137 1.0473 ± 0.0203 0.8457 ± 0.0031 0.7017 ± 0.0117 0.0149 ± 0.0055 0.5856 ± 0.0113 1.4664 ± 0.0237 Table 2. Point prediction and uncertainty evaluation on BDB2020+ independent test set. Results averaged over 20 seeds (± standard deviation). Best results in bold, second best underlined among aggregation methods. Method MAE ā RMSE ā Corr ā R^2 ā ECE ā CRPS ā NLL ā Individual Scoring Engines DynamicBind 0.8540 1.0860 0.4630 0.0110 - - - FlowDock 0.8630 1.1480 0.5290 -0.1040 - - - GNINA 0.7310 1.0450 0.4920 0.0840 - - - BIND 2.1000 2.3360 0.3100 -3.5750 - - - Traditional Aggregation Methods Consensus Scoring 0.7390 0.9970 0.5870 0.1660 - - - Ensemble Scoring 0.7990 1.0020 0.6200 0.1580 - - - LWE 0.7270 0.9620 0.5700 0.2240 - - - Softmax MoE 0.7560 ± 0.0400 0.9740 ± 0.0470 0.6150 ± 0.0240 0.2040 ± 0.0760 - - - Learned Aggregation Methods Baseline 0.7650 ± 0.0610 0.9900 ± 0.0770 0.5860 ± 0.0240 0.1750 ± 0.1340 - - - Gaussian 0.7400 ± 0.0470 0.9510 ± 0.0540 0.5920 ± 0.0230 0.2390 ± 0.0900 0.0820 ± 0.0280 0.5350 ± 0.0280 1.4150 ± 0.0430 MC Dropout 0.8100 ± 0.0570 1.0450 ± 0.0640 0.5590 ± 0.0310 0.0820 ± 0.1140 0.1430 ± 0.0250 0.6200 ± 0.0310 1.6160 ± 0.0420 Deep Ensemble 0.7230 ± 0.0350 0.9310 ± 0.0420 0.5950 ± 0.0210 0.2720 ± 0.0660 0.1040 ± 0.0240 0.5300 ± 0.0190 1.4240 ± 0.0290 SVGP 0.7730 ± 0.0580 1.0040 ± 0.0590 0.5550 ± 0.0450 0.1520 ± 0.0990 0.1430 ± 0.0410 0.5990 ± 0.0320 1.5910 ± 0.0620 SWAG 0.7480 ± 0.0430 0.9670 ± 0.0590 0.5860 ± 0.0370 0.2140 ± 0.0970 0.0830 ± 0.0260 0.5420 ± 0.0290 1.4300 ± 0.0430 Evidential Regression 0.7570 ± 0.0610 0.9680 ± 0.0740 0.5840 ± 0.0250 0.2110 ± 0.1310 0.0700 ± 0.0360 0.5440 ± 0.0420 1.4160 ± 0.0710 Ours RELIABLE-BA 0.7220 ± 0.0220 0.9320 ± 0.0180 0.5870 ± 0.0050 0.2720 ± 0.0280 0.0550 ± 0.0160 0.5150 ± 0.0100 1.3750 ± 0.0170 4.2.1. RQ1: Point Prediction Performance Table 1 reports point prediction performance on the PDBbind time-split test set. Among learned and traditional aggregation methods, RELIABLE-BA achieves the strongest overall performance, attaining the lowest MAE (0.8247), lowest RMSE (1.0473), highest correlation (0.8457), and highest R2R^2 (0.7017). RELIABLE-BA is also comparable to the strongest individual engine (DynamicBind: MAE 0.8095, RMSE: 1.0840) while providing calibrated uncertainty estimates that individual engines lack. Table 2 reports results on BDB2020+, an independent test set with zero overlap with training data. RELIABLE-BA achieves the best MAE (0.722) and R2R^2 (0.272), with competitive RMSE (0.932, second to Deep Ensembleās 0.931) again showing its high point prediction performance. 4.2.2. RQ2: Uncertainty Calibration and Quality Table 1 summarizes the uncertainty evaluation results on PDBbind time-split test set. RELIABLE-BA consistently outperforms all uncertainty-aware baselines in calibration and distributional quality, achieving the lowest ECE, CRPS, and NLL. Compared to evidential regression, RELIABLE-BA reduces ECE by 71.2%, CRPS by 6.0%, and NLL by 4.7%, indicating substantially improved calibration and sharper predictive distributions. Table 2 with BDB2020+ also shows that RELIABLE-BA consistently outperforms all uncertainty-aware baselines, again achieving the lowest ECE, CRPS, and NLL. Compared to evidential regression, RELIABLE-BA reduces ECE by 21.4%, CRPS by 5.3%, and NLL by 2.9%, confirming that the calibration improvements transfer to an independent validation set. 4.2.3. RQ3: Selective Prediction and RiskāCoverage Tradeoff Figure 3. Selective predictive analysis on five models. Selective predictive analysis on five models. We evaluate whether RELIABLE-BAās uncertainty estimates meaningfully correlate with prediction difficulty using selective prediction. Figure 3 reports relative MAE improvement as increasingly uncertain samples are discarded. RELIABLE-BA achieves consistent improvement across all discard thresholds, reaching 25.7% relative error reduction when retaining only the most confident 10% of predictions. RELIABLE-BA not only produces well-calibrated uncertainty estimates, but also ranks samples by prediction difficulty more effectively, enabling principled risk-aware filtering in virtual screening pipelines. 4.2.4. RQ4: Reliability Modeling and Expert Disagreement Figure 4. Correlation between expert disagreement and epistemic uncertainty of RELIABLE-BA (left) and expert disagreement and prediction error of RELIABLE-BA (right). Correlation between expert disagreement and epistemic uncertainty of RELIABLE-BA (left) and expert disagreement and prediction error of RELIABLE-BA (right). We analyze how RELIABLE-BAās learned reliability interacts with expert disagreement. Figure 4 (left) shows a clear positive correlation between inter-expert disagreement and RELIABLE-BAās epistemic uncertainty. As expert predictions diverge, RELIABLE-BA appropriately assigns higher epistemic uncertainty, treating disagreement as a signal of epistemic risk rather than noise. Figure 4 (right) shows that prediction error itself is only weakly correlated with expert disagreement, indicating that RELIABLE-BA does not force consensus at the level of the predictive mean. Instead, disagreement is absorbed through reliability-aware scaling of epistemic uncertainty, preserving stable point predictions while accurately reflecting uncertainty. This design enables RELIABLE-BA to remain relatively robust under heterogeneous expert behavior. 4.3. Ablation Studies We conduct ablation studies to isolate the contribution of key components in RELIABLE-BA, with a particular focus on reliability modeling and expert diversity. Specifically, we evaluate five ablations: (i) removing reliability scaling (No Reliability Scaling), (i) using only engine scores for reliability estimation without embeddings (Scores Only Reliability), (i) using uniform reliability weights instead of context-dependent reliability (Uniform Reliability), (iv) aggregating expert predictions using uniform weights without learned reliability-based weighting (Uniform Weight Aggregation), and (v) varying the number of docking engines used in the fusion. For engine diversity experiments, we select engines based on their reliability distribution: 2 engines (DynamicBind, FlowDock) and 3 engines (DynamicBind, FlowDock, BIND). Detailed selection choices are shown in Appendix E. Table 3. Ablation study on RELIABLE-BA components. Results averaged over 20 seeds. Method MAE ā RMSE ā Corr ā R^2 ā ECE ā CRPS ā NLL ā RELIABLE-BA 0.8247 1.0473 0.8457 0.7017 0.0149 0.5856 1.4664 No Reliability Scaling 0.8356 1.0613 0.8410 0.6937 0.0147 0.5939 1.4771 Scores Only Reliability 0.8385 1.0679 0.8414 0.6899 0.0146 0.5972 1.4851 Uniform Weight Aggregation 0.8383 1.0555 0.8394 0.6972 0.0209 0.5950 1.4796 Uniform Reliability 0.8418 1.0701 0.8404 0.6886 0.0279 0.5997 1.5004 As shown in Table 3, RELIABLE-BA achieves the best point prediction performance across most ablation variants, with the lowest MAE (0.8247) and RMSE (1.0473) and highest correlation (0.8457) and R2R^2 (0.7017). The primary benefit of reliability modeling is further reflected in uncertainty quality: removing molecular context from reliability estimation or using uniform weights consistently degrades calibration. While ECE is comparable across variants, RELIABLE-BA achieves the best CRPS (0.5856) and NLL (1.4664), demonstrating that context-dependent reliability modeling improves both predictive accuracy and uncertainty calibration. Table 4. Effect of the number of docking engines in RELIABLE-BA. Results averaged over 20 seeds. Method MAE ā RMSE ā Corr ā R^2 ā ECE ā CRPS ā NLL ā 4 Engines 0.8247 1.0473 0.8457 0.7017 0.0149 0.5856 1.4664 3 Engines 0.8471 1.1084 0.8448 0.6657 0.0349 0.6124 1.5190 2 Engines 0.8719 1.0889 0.8340 0.6775 0.0792 0.6245 1.5621 Table 4 shows that incorporating additional engines improves both accuracy and uncertainty quality. Adding a fourth engine reduces MAE by 5.4% relative to two engines and by 2.6% relative to three engines. Calibration benefits are more pronounced: ECE decreases by 81.2% compared to two engines and by 57.3% compared to three engines. These results confirm that RELIABLE-BA effectively leverages the complementary strengths of heterogeneous predictors, with each additional engine providing gains in prediction accuracy and calibration. 4.4. Case Study I: SARS-CoV-2 Mpro Dataset We have conducted case study evaluation on the SARS-CoV-2 main protease dataset (Mpro), an independent real-world benchmark disjoint from our training data (24). Evaluating on the SARS-CoV-2 main protease dataset, a target central to the COVID-19 pandemic response, demonstrates the real-world impact of this study. RELIABLE-BA outperforms all individual engines in terms of point prediction with best MAE (0.532±0.0210.532± 0.021), RMSE (0.644±0.0240.644± 0.024) and correlation (0.695±0.0060.695± 0.006) compared to individual engines like GNINA (MAE: 0.631, RMSE: 0.768, Correlation: 0.586) and DynamicBind (MAE: 0.643, RMSE: 0.780, Correlation: 0.683). RELIABLE-BA also showed strong uncertainty calibration (CRPS = 0.391, NLL = 1.144), demonstrating effective generalization to systematic out-of-domain evaluation. 4.5. Case Study I: 5HT2A Receptor Figure 5. Experimentally resolved structure of the human 5HT2A receptor (PDB ID: 7WC6) used in the case study evaluation. Structure of human 5HT2A receptor. The 5HT2A receptor is a class A G proteinācoupled receptor (GPCR) that plays a central role in serotonergic signaling and is a key therapeutic target for neurological and psychiatric disorders, including depression and anxiety (Figure 5) (37). From a computational perspective, 5HT2A provides a realistic evaluation setting for affinity prediction under distribution shift, as the resolved receptor structure is not represented in the training split and the ligand set exhibits substantial functional heterogeneity. This setting is well-suited for examining the behavior of uncertainty-aware fusion methods, and thus can prove that this study can expand to virtual-screening settings. We select PDB structure 7WC6 as the protein target and evaluate all methods on distinct ligands, averaged over 20 random seeds. RELIABLE-BA achieves an MAE of 1.002±0.041.002± 0.04 and an RMSE of 1.264±0.0431.264± 0.043, outperforming the strongest individual engine, DynamicBind (MAE: 1.060, RMSE: 1.360), by 5.5% and 7.1%, respectively. Notably, uncertainty calibration remains strong under distribution shift. RELIABLE-BA achieves ECE of 0.0233, comparable to in-distribution performance (0.0149), with CRPS of 0.7143 and NLL of 1.6622, showing competitive uncertainty quantification. 5. Conclusion In this paper, we proposed RELIABLE-BA, a reliability-aware evidential framework that fuses heterogeneous docking engines through context-dependent uncertainty modulation. By modeling each engine as an evidential expert and scaling epistemic uncertainty according to learned reliability, RELIABLE-BA produces calibrated predictive distributions that enable effective selective prediction, reducing MAE by 25.7% when retaining the most confident 10% of predictions. Experiments on PDBbind and BDB2020+ demonstrate strong point prediction accuracy and well-calibrated uncertainty estimates, with additional validation on the SARS-CoV-2 Mpro dataset and the 5HT2A GPCR confirming applicability to practical drug targets. Moreover, RELIABLE-BAās modular design allows it to integrate heterogeneous docking and scoring engines, making it broadly applicable to large-scale virtual screening pipelines. Future work will explore extending RELIABLE-BA to larger docking ensembles and incorporating pose-level uncertainty to further enhance applicability in drug discovery workflows. 6. Limitations and Ethical Considerations This work does not involve human participants or personally identifiable information. All experiments are conducted on publicly available benchmark datasets for proteināligand binding affinity prediction (PDBbind 2020 (43), ChEMBL (13), BDB2020+ (24)), derived from previously published experimental measurements and used in accordance with their original licenses. As a limitation, RELIABLE-BA relies on the availability of multiple binding affinity prediction engines and may be less effective when only a few engines are available or when their predictions are highly correlated. The learned reliability scores reflect context-dependent predictive trustworthiness rather than physical or mechanistic correctness and should not be interpreted as causal guarantees. Our evaluation relies on established benchmark datasets, which may contain inherent biases or data leakage despite mitigation efforts. Finally, RELIABLE-BA introduces computational overhead from multi-engine inference, though this overhead is minimal compared to the cost of individual docking engines (Appendix F). Acknowledgements.This work was supported in part by the U.S. National Science Foundation under Grant No. 2125142 to Y.L. and by the National Institutes of Health under Grant No. R35GM153437 to V.K. GPU resources were provided through an NVIDIA Academic Grant. GenAI Disclosure Generative AI tools were used in a limited capacity. Claude was utilized for minor code construction, while ChatGPT was used for grammar and writing refinement. References [1] A. Amini, W. Schwarting, A. Soleimany, and D. Rus (2020) Deep evidential regression. Advances in neural information processing systems 33, p. 14927ā14937. Cited by: §1, §2.3, §3.2, §4.1.3. [2] P. J. Ballester and J. B. O. Mitchell (2010) A machine learning approach to predicting proteināligand binding affinity with applications to molecular docking. Bioinformatics 26 (9), p. 1169ā1175. External Links: Document Cited by: §1, §2.2. [3] C. Blanes-Mira, P. FernĆ”ndez-Aguado, J. de AndrĆ©s-López, A. FernĆ”ndez-Carvajal, A. Ferrer-Montiel, and G. FernĆ”ndez-Ballester (2022) Comprehensive survey of consensus docking for high-throughput virtual screening. Molecules 28 (1), p. 175. External Links: Document Cited by: §1, §2.2. [4] L. Breiman (1996) Stacked regressions. Machine Learning 24 (1), p. 49ā64. External Links: Document Cited by: Appendix C. [5] M. Buttenschoen, G. M. Morris, and C. M. Deane (2024) PoseBusters: ai-based docking methods fail to generate physically valid poses or generalise to novel sequences. Chemical Science 15 (9), p. 3130ā3139. Cited by: §1. [6] D. Cao, J. Enouen, Y. Wang, X. Song, C. Meng, H. Niu, and Y. Liu (2023) Estimating treatment effects from irregular time series observations with hidden confounders. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, p. 6897ā6905. Cited by: §1. [7] D. Cao, W. Ye, Y. Zhang, S. Griesemer, and Y. Liu (2026) PINFDit: energy-based physics-informed Diffusion Transformers for general-purpose time series tasks. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1. [8] S. Chithrananda, G. Grand, and B. Ramsundar (2020) ChemBERTa: large-scale self-supervised pretraining for molecular property prediction. arXiv preprint arXiv:2010.09885. Cited by: §3.1, §4.1.1. [9] Y. Chung, I. Char, H. Guo, J. Schneider, and W. Neiswanger (2021) Uncertainty toolbox: an open-source library for assessing, visualizing, and improving uncertainty quantification. arXiv preprint arXiv:2109.10254. Cited by: §4.1.4. [10] S. S. Ericksen, H. Wu, H. Zhang, L. A. Michael, M. A. Newton, F. M. Hoffmann, and S. A. Wildman (2017) Machine learning consensus scoring improves performance across targets in structure-based virtual screening. Journal of Chemical Information and Modeling 57 (7), p. 1579ā1590. External Links: Document Cited by: §1, §2.2. [11] R. A. Friesner, J. L. Banks, R. B. Murphy, T. A. Halgren, J. J. Klicic, D. T. Mainz, M. P. Repasky, E. H. Knoll, M. Shelley, J. K. Perry, et al. (2004) Glide: a new approach for rapid, accurate docking and scoring. 1. method and assessment of docking accuracy. Journal of Medicinal Chemistry 47 (7), p. 1739ā1749. Cited by: §2.1. [12] Y. Gal and Z. Ghahramani (2016) Dropout as a bayesian approximation: representing model uncertainty in deep learning. In international conference on machine learning, p. 1050ā1059. Cited by: §2.3, §4.1.3. [13] A. Gaulton, L. J. Bellis, A. P. Bento, J. Chambers, M. Davies, A. Hersey, Y. Light, S. McGlinchey, D. Michalovich, B. Al-Lazikani, et al. (2012) ChEMBL: a large-scale bioactivity database for drug discovery. Nucleic acids research 40 (D1), p. D1100āD1107. Cited by: §6. [14] D. E. Graff, E. I. Shakhnovich, and C. W. Coley (2021) Accelerating high-throughput virtual screening through molecular pool-based active learning. Chemical science 12 (22), p. 7866ā7881. Cited by: §1. [15] A. S. Hauser, M. M. Attwood, M. Rask-Andersen, H. B. Schiƶth, and D. E. Gloriam (2017) Trends in gpcr drug discovery: new agents, targets and indications. Nature reviews Drug discovery 16 (12), p. 829ā842. Cited by: §1. [16] J. Hensman, A. Matthews, and Z. Ghahramani (2015) Scalable variational gaussian process classification. In Artificial intelligence and statistics, p. 351ā360. Cited by: §2.3, §4.1.3. [17] J. JimĆ©nez, M. Skalic, G. Martinez-Rosell, and G. De Fabritiis (2018) K deep: proteināligand absolute binding affinity prediction via 3d-convolutional neural networks. Journal of chemical information and modeling 58 (2), p. 287ā296. Cited by: §2.1. [18] T. Kalliokoski, C. Kramer, A. Vulpetti, and P. Gedeck (2013) Comparability of mixed ic50 dataāa statistical analysis. PLoS ONE 8 (4), p. e61007. External Links: Document Cited by: §4.1.1. [19] A. Kendall and Y. Gal (2017) What uncertainties do we need in bayesian deep learning for computer vision?. In Advances in Neural Information Processing Systems, Vol. 30. Cited by: §1, §2.3, §4.1.3. [20] B. Lakshminarayanan, A. Pritzel, and C. Blundell (2017) Simple and scalable predictive uncertainty estimation using deep ensembles. Vol. 30. Cited by: §2.3, §4.1.3. [21] H. Y. I. Lam, J. S. Guan, X. E. Ong, R. Pincket, and Y. Mu (2024) Protein language models are performant in structure-free virtual screening. Briefings in Bioinformatics 25 (6), p. bbae480. Cited by: §1, §2.1, §2.1, 2nd item. [22] J. H. Lam and V. Katritch (2025) Navigating structure-based drug discovery with emerging innovations in physics-and knowledge-based approaches. npj Drug Discovery 2 (1), p. 29. Cited by: §1, §2.1. [23] G. A. Landrum and S. Riniker (2024) Combining ic50 or k i values from different sources is a source of significant noise. Journal of chemical information and modeling 64 (5), p. 1560ā1567. Cited by: §2.1, §4.1.1. [24] J. Li, X. Guan, O. Zhang, K. Sun, Y. Wang, D. Bagni, and T. Head-Gordon (2026) Leak proof pdbbind: a reorganized data set of proteināligand complexes for more generalizable binding affinity prediction. The Journal of Physical Chemistry B 130 (2), p. 730ā740. Cited by: Appendix B, §1, §1, §4.1.1, §4.4, §6. [25] Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhu, W. Lu, A. dos Santos Costa, M. Fazel-Zarandi, T. Sercu, S. Candido, et al. (2022) Language models of protein sequences at the scale of evolution enable accurate structure prediction. BioRxiv 2022, p. 500902. Cited by: §3.1, §4.1.1. [26] W. Lu, Q. Wu, J. Zhang, J. Rao, C. Li, and S. Zheng (2024) DynamicBind: predicting ligand-specific protein-ligand complex structure with a deep equivariant generative model. Nature Communications 15 (1), p. 1071. Cited by: §1, §2.1, 4th item, §4.1.1. [27] H. Ma, Z. Han, C. Zhang, H. Fu, J. T. Zhou, and Q. Hu (2021) Trustworthy multimodal regression with mixture of normal-inverse gamma distributions. Advances in Neural Information Processing Systems 34, p. 6881ā6893. Cited by: §A.4, §1, §2.3, §3.4. [28] W. J. Maddox, P. Izmailov, T. Garipov, D. P. Vetrov, and A. G. Wilson (2019) A simple baseline for bayesian uncertainty in deep learning. Vol. 32. Cited by: §2.3, §4.1.3. [29] A. T. McNutt, P. Francoeur, R. Aggarwal, T. Masuda, R. Meli, M. Ragoza, J. Sunseri, and D. R. Koes (2021) GNINA 1.0: molecular docking with deep learning. Journal of cheminformatics 13 (1), p. 43. Cited by: §1, §2.1, 1st item. [30] J. Mohamed Abdul Cader, M. H. Newton, J. Rahman, A. J. Mohamed Abdul Cader, and A. Sattar (2024) Ensembling methods for protein-ligand binding affinity prediction. Scientific Reports 14 (1), p. 24447. Cited by: §1, §2.2. [31] A. Morehead and J. Cheng (2025) FlowDock: geometric flow matching for generative proteināligand docking and affinity prediction. Bioinformatics 41 (Supplement_1), p. i198āi206. External Links: Document Cited by: §2.1, 3rd item, §4.1.1. [32] D. Nhat Phuong, D. R. Flower, S. Chattopadhyay, and A. K. Chattopadhyay (2023) Towards effective consensus scoring in structure-based virtual screening. Interdisciplinary Sciences: Computational Life Sciences 15 (1), p. 131ā145. External Links: Document Cited by: §1, §2.2. [33] F. Panahandeh and N. Mansouri (2025) A comprehensive review of neural network-based approaches for drugātarget interaction prediction. Molecular Diversity, p. 1ā48. Cited by: §1. [34] S. Passaro, G. Corso, J. Wohlwend, M. Reveiz, S. Thaler, V. R. Somnath, N. Getz, T. Portnoi, J. Roy, H. Stark, et al. (2025) Boltz-2: towards accurate and efficient binding affinity prediction. bioRxiv. External Links: Document Cited by: §1. [35] M. Rayka, M. Mirzaei, and A. Mohammad Latifi (2024) An ensemble-based approach to estimate confidence of predicted proteināligand binding affinity values. Molecular Informatics 43 (4), p. e202300292. External Links: Document Cited by: §1, §2.2. [36] M. Rayka and S. S. Naghavi (2025) Uncertainty quantification enables reliable deep learning for proteināligand binding affinity prediction. Scientific Reports 15 (1), p. 43156. External Links: Document Cited by: §1, §2.2. [37] C. J. Schmidt, S. M. Sorensen, J. H. Kenne, A. A. Carr, and M. G. Palfreyman (1995) The role of 5-ht2a receptors in antipsychotic activity. Life sciences 56 (25), p. 2209ā2222. Cited by: §4.5. [38] A. P. Soleimany, A. Amini, S. Goldman, D. Rus, S. N. Bhatia, and C. W. Coley (2021) Evidential deep learning for guided molecular property prediction and discovery. ACS central science 7 (8), p. 1356ā1367. External Links: Document Cited by: §2.2. [39] H. StƤrk, O. Ganea, L. Pattanaik, R. Barzilay, and T. Jaakkola (2022) Equibind: geometric deep learning for drug binding structure prediction. In International conference on machine learning, p. 20503ā20521. Cited by: §4.1.1. [40] M. M. Stepniewska-Dziubinska, P. Zielenkiewicz, and P. Siedlecki (2018) Development and evaluation of a deep learning model for proteināligand binding affinity prediction. Bioinformatics 34 (21), p. 3666ā3674. Cited by: §2.1. [41] O. Trott and A. J. Olson (2010) AutoDock vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. Journal of Computational Chemistry 31 (2), p. 455ā461. Cited by: §2.1. [42] L. Wang, Y. Wu, Y. Deng, B. Kim, L. Pierce, G. Krilov, D. Lupyan, S. Robinson, M. K. Dahlgren, J. Greenwood, et al. (2015) Accurate and reliable prediction of relative ligand binding potency in prospective drug discovery by way of a modern free-energy calculation protocol and force field. Journal of the American Chemical Society 137 (7), p. 2695ā2703. Cited by: §1, §2.1. [43] R. Wang, X. Fang, Y. Lu, C. Yang, and S. Wang (2005) The pdbbind database: methodologies and updates. Journal of medicinal chemistry 48 (12), p. 4111ā4119. Cited by: §1, §4.1.1, §6. [44] Z. Wang, H. Sun, X. Yao, D. Li, L. Xu, Y. Li, S. Tian, and T. Hou (2016) Comprehensive evaluation of ten docking programs on a diverse set of proteināligand complexes: the prediction accuracy of sampling power and scoring power. Physical Chemistry Chemical Physics 18 (18), p. 12964ā12975. External Links: Document Cited by: §1, §2.2. [45] W. Xu, X. Liu, J. Wang, F. Zhang, D. Hu, and L. Zong (2025) UAMRL: multi-granularity uncertainty-aware multimodal representation learning for drug-target affinity prediction. Bioinformatics 41 (10), p. btaf512. External Links: Document Cited by: §1, §2.2. [46] J. Yang, Y. Chen, T. Shen, B. S. Kristal, and D. F. Hsu (2005) Consensus scoring criteria for improving enrichment in virtual screening. Journal of Chemical Information and Modeling 45 (4), p. 1134ā1146. External Links: Document Cited by: §1, §2.2. [47] Y. Zhao, Y. Xing, Y. Zhang, Y. Wang, M. Wan, D. Yi, C. Wu, S. Li, H. Xu, H. Zhang, et al. (2025) Evidential deep learning-based drug-target interaction prediction. Nature communications 16 (1), p. 6915. Cited by: §1, §2.2. Appendix A Proof of Validity of Reliability-Scaled MoNIG Aggregation We prove that RELIABLE-BAās reliability scaling preserves NIG validity and that MoNIG fusion remains valid under this scaling. A.1. Preliminaries A NormalāInverse-Gamma distribution NIGā(γ,ν,α,β)NIG(γ,ν,α,β) is valid if: (V1) ν>0ν>0 (V2) α>1α>1 (V3) β>0β>0 Under these conditions, the moments are well-defined: (21) ā[μ] [μ] =γ, =γ, (22) ā[Ļ2] [Ļ^2] =βαā1(aleatoric uncertainty), = βα-1 (aleatoric uncertainty), (23) Varā(μ) (μ) =βνā(αā1)(epistemic uncertainty). = βν(α-1) (epistemic uncertainty). A.2. Reliability Scaling Preserves Validity Given valid NIG parameters (γj,νj,αj,βj)( _j, _j, _j, _j) and reliability score rjā(0,1]r_jā(0,1], RELIABLE-BA applies the following transformation: (24) γ~j γ_j =γj, = _j, (25) ν~j ν_j =rjāνj, =r_j _j, (26) α~j α_j =1+rjā(αjā1), =1+r_j( _j-1), (27) β~j β_j =rjāβj. =r_j _j. Proposition A.1. If (γj,νj,αj,βj)( _j, _j, _j, _j) is a valid NIG and rjā(0,1]r_jā(0,1], then (γ~j,ν~j,α~j,β~j)( γ_j, ν_j, α_j, β_j) is also valid. Proof. We verify each validity condition: (V1) ν~j=rjāνj>0 ν_j=r_j _j>0 since rj>0r_j>0 and νj>0 _j>0. (V2) α~j=1+rjā(αjā1)>1 α_j=1+r_j( _j-1)>1 since rj>0r_j>0 and αjā1>0 _j-1>0. (V3) β~j=rjāβj>0 β_j=r_j _j>0 since rj>0r_j>0 and βj>0 _j>0. ā A.3. Effect on Uncertainty Components Proposition A.2 (Aleatoric Invariance). Reliability scaling preserves aleatoric uncertainty: Ļ~a,j2=Ļa,j2 Ļ^2_a,j=Ļ^2_a,j. Proof. (28) Ļ~a,j2=β~jα~jā1=rjāβjrjā(αjā1)=βjαjā1=Ļa,j2. Ļ^2_a,j= β_j α_j-1= r_j _jr_j( _j-1)= _j _j-1=Ļ^2_a,j. ā Proposition A.3 (Epistemic Scaling). Reliability scaling increases epistemic uncertainty for unreliable experts: Ļ~e,j2=Ļe,j2/rj Ļ^2_e,j=Ļ^2_e,j/r_j. Proof. (29) Ļ~e,j2=β~jν~jā(α~jā1)=rjāβjrjāνjā rjā(αjā1)=1rjā βjνjā(αjā1)=Ļe,j2rj. Ļ^2_e,j= β_j ν_j( α_j-1)= r_j _jr_j _jĀ· r_j( _j-1)= 1r_jĀ· _j _j( _j-1)= Ļ^2_e,jr_j. ā When rj<1r_j<1, epistemic uncertainty increases by factor 1/rj>11/r_j>1, reflecting reduced confidence in unreliable experts. A.4. MoNIG Fusion Under Reliability Scaling Following Ma et al. (27), the NIG summation operator ā fuses M distributions with parameters: (30) γf _f =āj=1Mνjāγjāj=1Mνj, = _j=1^M _j _j _j=1^M _j, (31) νf _f =āj=1Mνj, = _j=1^M _j, (32) αf _f =āj=1Mαj+Mā12, = _j=1^M _j+ M-12, (33) βf _f =āj=1Mβj+12āāj=1Mνjā(γjāγf)2. = _j=1^M _j+ 12 _j=1^M _j( _j- _f)^2. Theorem A.4 (Validity of Reliability-Scaled Fusion). Let (γj,νj,αj,βj)j=1M\( _j, _j, _j, _j)\_j=1^M be M valid NIG distributions with reliability scores rjā(0,1]r_jā(0,1]. Then the fused distribution (34) āØj=1MNIGā(γ~j,ν~j,α~j,β~j) _j=1^MNIG( γ_j, ν_j, α_j, β_j) is a valid NIG. Proof. We verify each validity condition for the fused parameters. Condition (V1): (35) νf=āj=1Mν~j=āj=1Mrjāνj>0, _f= _j=1^M ν_j= _j=1^Mr_j _j>0, since each term rjāνj>0r_j _j>0. Condition (V2): Substituting the scaled α~j α_j: (36) αf _f =āj=1Mα~j+Mā12 = _j=1^M α_j+ M-12 (37) =āj=1M(1+rjā(αjā1))+Mā12 = _j=1^M (1+r_j( _j-1) )+ M-12 (38) =M+āj=1Mrjā(αjā1)+Mā12 =M+ _j=1^Mr_j( _j-1)+ M-12 (39) =3āMā12+āj=1Mrjā(αjā1). = 3M-12+ _j=1^Mr_j( _j-1). Since rj>0r_j>0 and αj>1 _j>1 for all j, we have αf>3āMā12ā„1 _f> 3M-12ā„ 1 for Mā„1Mā„ 1. Condition (V3): The fused scale parameter is: (40) βf=āj=1Mrjāβjā>0+12āāj=1Mν~jā(γ~jāγf)2āā„0>0. _f= _j=1^Mr_j _j_>0+ 12 _j=1^M ν_j( γ_j- _f)^2_ā„ 0>0. ā Corollary A.5 (Reliability-Weighted Fusion Mean). Under reliability scaling, the fused mean becomes: (41) γf=āj=1Mrjāνjāγjāj=1Mrjāνj. _f= _j=1^Mr_j _j _j _j=1^Mr_j _j. Each expertās contribution is weighted by rjāνjr_j _j, the product of reliability and evidence. A.5. Summary The reliability scaling mechanism in RELIABLE-BA satisfies the following properties: (1) Scaled parameters remain valid NIG distributions (Proposition 1). (2) Aleatoric uncertainty is preserved exactly (Proposition 2). (3) Epistemic uncertainty scales inversely with reliability (Proposition 3). (4) MoNIG fusion of reliability-scaled experts yields a valid NIG (Theorem 1). (5) The fused mean weights experts by reliability Ć evidence (Corollary 1). Appendix B Dataset Curation For PDBbind, after filtering invalid entries and complexes with incomplete predictions from engines, the final data consists of 8,742 unique proteināligand complexes (7,924 train / 576 validation / 242 test). For BDB2020+, we retained 96 complexes after excluding entries with proteins exceeding embedding length limits or incomplete structural annotations. For the 5HT2A case study, we evaluate on 3,868 ligands with valid affinity annotations and docking engine outputs curated from ChEMBL dataset. For SARS-CoV-2 Mpro dataset, we utilized the curated dataset in LP-PDBBind (24). Appendix C Baseline Architectures and Hyperparameters All neural network baselines use the same encoder: a 4-layer network (708 ā 512 ā 256 ā 128 ā 64) that takes 4 engine scores and 704-dimensional molecular embeddings as input. We use ReLU activations, 20% dropout, Adam optimizer with learning rate 0.0005, batch size 64, and stop training if validation MAE doesnāt improve for 30 epochs. RELIABLE-BA Each docking engine has its own 3-layer head (1 ā 256 ā 128 ā 64) that maps the engine score to four NIG parameters through separate linear projections. A separate reliability network (704 ā 128 ā 4) takes molecular embeddings and outputs a reliability score for each engine via sigmoid activation. Both components train jointly with evidential loss (Ī»=0.001Ī»=0.001). Evidential Regression Single evidential network that predicts four NIG parameters (γ,ν,α,βγ,ν,α,β) from all inputs combined. Uses shared encoder with 4 linear output heads. Same evidential loss (Ī»=0.001Ī»=0.001) but without per-engine modeling or reliability weighting. Gaussian Predicts mean μ and variance Ļ2Ļ^2 directly. Uses shared encoder with 2 linear output heads. Trained with Gaussian negative log-likelihood. Baseline Standard regression network without uncertainty estimation. Uses shared encoder with single linear output. Trained with mean squared error loss. Deep Ensemble MVE Five independently trained Gaussian networks with different random initializations. Final prediction averages across members. MC Dropout Gaussian network with dropout kept active during prediction. Runs 50 forward passes and estimates uncertainty from variation across passes. SVGP Gaussian process with learned deep features. Encoder (without dropout) maps to 64 dimensions, followed by RBF kernel GP with 128 inducing points. SWAG Gaussian base model that collects weight statistics after epoch 75. At prediction, samples 30 weight configurations from the approximate posterior. Consensus Scoring Computes median across engines, then weights each engine by expā”(ā|deviation from median|) (-|deviation from median|). Ensemble Scoring Arithmetic mean of engines with standard deviation across engines. LWE Linear stacking (4) with four fixed softmax-normalized scalar weights on raw scores, plus one learnable noise scalar for uncertainty. Softmax MoE Mixture-of-experts on the full input. Four expert networks (708 ā 128 ā 64 ā 1) combined by an input-dependent softmax gating network, with a separate network for uncertainty. Appendix D Hyperparameter Sensitivity Table 5 shows RELIABLE-BA performance across different values of the evidential regularization coefficient Ī». Lower Ī» yields substantially better calibration (ECE, NLL), while moderate values (Ī»=0.01Ī»=0.01) slightly improve point prediction at the expense of calibration. We select Ī»=0.001Ī»=0.001 as it achieves the best calibration with only a marginal increase in MAE, reflecting our emphasis on well-calibrated uncertainty estimates. Table 5. Effect of evidential regularization coefficient Ī» on RELIABLE-BA performance. Results averaged over 20 seeds (std in parentheses). Ī» MAE ā RMSE ā ECE ā NLL ā 0.001 0.824 (0.015) 1.045 (0.021) 0.014 (0.003) 1.463 (0.023) 0.005 0.816 (0.019) 1.039 (0.024) 0.027 (0.011) 1.474 (0.030) 0.01 0.814 (0.017) 1.037 (0.022) 0.042 (0.011) 1.491 (0.032) 0.05 0.823 (0.041) 1.046 (0.051) 0.078 (0.019) 1.596 (0.086) Appendix E Reliability Score Visualization Figure 6. t-SNE visualization of protein-ligand embeddings colored by learned reliability scores. t-SNE of protein-ligand embeddings by reliability scores. Figure 6 visualizes learned reliability scores across proteināligand embedding space using t-SNE on the PDBBind test set. Each panel shows the reliability assigned to one docking engine, with color indicating reliability (yellow = high, purple = low). FlowDock achieves the highest mean reliability (0.895), followed by DynamicBind (0.81), GNINA (0.719), and BIND (0.706). Appendix F Computational Cost Analysis Table 6 reports training and inference times for all methods on the PDBBind benchmark (7,924 training / 242 test samples). All experiments were conducted on a single NVIDIA RTX 5090 GPU. Table 6. Computational cost comparison. Training and inference times in seconds, averaged over 20 runs. Method Training (s) Inference (s) Baseline 9.02 ± 0.64 2.04 ± 0.00 Gaussian 10.30 ± 0.54 2.72 ± 0.02 SWAG 10.19 ± 0.28 2.73 ± 0.01 MC Dropout 13.31 ± 0.94 2.77 ± 0.01 Evidential Regression 15.38 ± 1.06 2.74 ± 0.02 SVGP 27.85 ± 2.59 2.80 ± 0.01 Deep Ensemble MVE 30.14 ± 1.24 2.76 ± 0.01 RELIABLE-BA 48.38 ± 4.70 2.77 ± 0.01 RELIABLE-BA has higher total training time (48s) due to per-engine modeling, but training completes in under one minute and inference (11ms per complex) matches all other methods. In practice, upstream docking engines dominate computational cost (seconds to minutes per complex), making RELIABLE-BAās overhead negligible.