Paper deep dive
Multi-Task Genetic Algorithm with Multi-Granularity Encoding for Protein-Nucleotide Binding Site Prediction
Yiming Gao, Liuyi Xu, Pengshan Cui, Yining Qian, An-Yang Lu, Xianpeng Wang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/22/2026, 5:14:40 AM
Summary
MTGA-MGE is a novel framework for protein-nucleotide binding site prediction that addresses feature redundancy and rigid fusion mechanisms. It utilizes a Multi-Granularity Encoding (MGE) network for refined feature extraction, a Multi-Task Genetic Algorithm (MTGA) based on NSGA-III for adaptive architecture search, and an External-Neighborhood Mechanism (ENM) to facilitate cross-task knowledge transfer.
Entities (6)
Relation Signals (5)
MTGA-MGE → includes → MGE
confidence 100% · The MTGA-MGE framework (Fig. 2) comprises three core components: Multi-Granularity Encoding (MGE)...
MTGA-MGE → includes → MTGA
confidence 100% · The MTGA-MGE framework (Fig. 2) comprises three core components: ... Multi-Task Genetic Algorithm (MTGA)...
MTGA-MGE → includes → ENM
confidence 100% · The MTGA-MGE framework (Fig. 2) comprises three core components: ... External-Neighborhood Mechanism (ENM).
MGE → uses → ProtT5
confidence 95% · we first employ ProtT5 to capture global contextual information.
MGE → uses → ESM2
confidence 95% · we utilize ESM2 to extract features rich in evolutionary signals.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Accurate identification of protein-nucleotide binding sites is fundamental to deciphering molecular mechanisms and accelerating drug discovery. However, current computational methods often struggle with suboptimal performance due to inadequate feature representation and rigid fusion mechanisms, which hinder the effective exploitation of cross-task information synergy. To bridge this gap, we propose MTGA-MGE, a framework that integrates a Multi-Task Genetic Algorithm with Multi-Granularity Encoding to enhance binding site prediction. Specifically, we develop a Multi-Granularity Encoding (MGE) network that synergizes multi-scale convolutions and self-attention mechanisms to distill discriminative signals from high-dimensional, redundant biological data. To overcome the constraints of static fusion, a genetic algorithm is employed to adaptively evolve task-specific fusion strategies, thereby effectively improving model generalization. Furthermore, to catalyze collaborative learning, we introduce an External-Neighborhood Mechanism (ENM) that leverages biological similarities to facilitate targeted information exchange across tasks. Extensive evaluations on fifteen nucleotide datasets demonstrate that MTGA-MGE not only establishes a new state-of-the-art in data-abundant, high-resource scenarios but also maintains a robust competitive edge in rare, low-resource regimes, presenting a highly adaptive scheme for decoding complex protein-ligand interactions in the post-genomic era.
Tags
Links
- Source: https://arxiv.org/abs/2603.14797v1
- Canonical: https://arxiv.org/abs/2603.14797v1
Trouble viewing inline? Open PDF directly →
Full Text
67,166 characters extracted from source content.
Expand or collapse full text
Multi-Task Genetic Algorithm with Multi-Granularity Encoding for Protein-Nucleotide Binding Site Prediction Yiming Gao, Liuyi Xu, Pengshan Cui, Yining Qian*, An-Yang Lu, Xianpeng Wang This work was supported in part by the National Natural Science Foundation of China (NSFC) under Grant Nos. 62522309, and in part by the Natural Science Foundation of Liaoning Province under Grant No. 2025JH6/101000012. (Corresponding author: Yining Qian.)This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.Yiming Gao, Liuyi Xu, Pengshan Cui and An-Yang Lu are with the College of Information Science and Engineering, Northeastern University, Shenyang, 110819, China (e-mail: gaoym3@mails.neu.edu.cn; xuliuyi@mails.neu.edu.cn; ccuipengshan@163.com; luanyang@mail.neu.edu.cn).Yining Qian is with the School of Computer Science and Engineering, Northeastern University, Shenyang 110819, China (e-mail: qianyiningning@126.com). Corresponding author.Xianpeng Wang is with the Key Laboratory of Data Analytics and Optimization for Smart Industry, Ministry of Education, Northeastern University, Shenyang, 110819, China (wangxianpeng@ise.neu.edu.cn). Abstract Accurate identification of protein-nucleotide binding sites is fundamental to deciphering molecular mechanisms and accelerating drug discovery. However, current computational methods often struggle with suboptimal performance due to inadequate feature representation and rigid fusion mechanisms, which hinder the effective exploitation of cross-task information synergy. To bridge this gap, we propose MTGA-MGE, a framework that integrates a Multi-Task Genetic Algorithm with Multi-Granularity Encoding to enhance binding site prediction. Specifically, we develop a Multi-Granularity Encoding (MGE) network that synergizes multi-scale convolutions and self-attention mechanisms to distill discriminative signals from high-dimensional, redundant biological data. To overcome the constraints of static fusion, a genetic algorithm is employed to adaptively evolve task-specific fusion strategies, thereby effectively improving model generalization. Furthermore, to catalyze collaborative learning, we introduce an External-Neighborhood Mechanism (ENM) that leverages biological similarities to facilitate targeted information exchange across tasks. Extensive evaluations on fifteen nucleotide datasets demonstrate that MTGA-MGE not only establishes a new state-of-the-art in data-abundant, high-resource scenarios but also maintains a robust competitive edge in rare, low-resource regimes, presenting a highly adaptive scheme for decoding complex protein-ligand interactions in the post-genomic era. I Introduction Protein-nucleotide binding site prediction is a cornerstone of molecular biology, aiming to pinpoint specific residues that interact with ligands such as DNA, RNA, and various nucleotides, thereby elucidating essential biological processes and guiding structural drug design [6, 42, 1, 2]. As visualized in Fig. 1, Figure 1: Schematic representation of protein–nucleotide interaction, further illustrating the action of a drug at the protein–nucleotide binding sites. The highlighted regions (orange/yellow) illustrate the binding residues distributed across the protein structure. the spatial arrangement of these residues defines the interaction interface, imposing strict constraints for protein function annotation and nucleotide-targeted drug discovery [3, 33, 30]. Without high-precision site identification, elucidating complex biological mechanisms at the molecular level remains elusive, while the lack of functional annotation hinders the efficiency of drug discovery [35]. Although experimental techniques such as X-ray crystallography, site-directed mutagenesis, and cross-linking provide reliable evidence [31, 4, 32], their high costs and low throughput render them inadequate for the exponentially growing volume of protein sequences. Consequently, data-driven computational approaches have emerged as the dominant research paradigm [10]. Current computational methods can be broadly categorized into nucleotide-specific models [7, 17, 14] and nucleotide-general models [12, 49, 48, 11, 40, 39]. Nucleotide-specific models, such as DeepGTP [17] and DeepATPseq [14], typically employ elaborate feature engineering for a specific ligand. While achieving acceptable accuracy on specific tasks, their critical limitation lies in their “isolated modeling” paradigm. By treating each nucleotide-binding event as an independent problem, these methods completely fail to leverage the shared biological homologies across different ligands [39]. Ignoring these cross-task homologies leads to an excessive reliance on task-specific data, restricting generalizability and resulting in suboptimal performance in low-resource scenarios where labeled samples for niche ligands are scarce [27]. To overcome this limitation, recent research has pivoted toward nucleotide-general models utilizing Multi-Task Learning (MTL). For instance, approaches like NucGMTL [40] and NucMoMTL [39] employ multi-task learning to jointly train on diverse nucleotide-binding tasks, leveraging shared biological rules to benefit low-resource ligands [38]. Furthermore, these methods increasingly rely on Pre-trained Protein Language Models (PLMs) [5, 13, 21] for sequence feature extraction. While PLMs provide rich global context, their extremely high-dimensional embeddings [34] are often processed using simple linear methods (e.g., mean pooling). This “coarse-grained” extraction obscures critical binding features with high-dimensional redundant noise [26], reducing sensitivity to weak local signals. Similarly, the simplistic concatenation of traditional handcrafted features also exacerbates dimensional redundancy and feature conflicts, further drowning out these weak signals. Because nucleotide binding is determined by subtle spatial arrangements of local residues, losing these faint signals prevents models from precisely pinpointing binding anchors within complex backgrounds [38]. Beyond feature-level representation, the structural integration of these features presents a further challenge. Existing nucleotide-general models predominantly adopt a shared-backbone fusion paradigm [28]. Although parameter sharing formally induces inter-task correlation, fixed, manually designed architectures fail to dynamically accommodate the biological heterogeneity across diverse nucleotide tasks [23]. Consequently, models are forced to compromise conflicting ligand demands through constrained weight fine-tuning. Mathematically, discovering task-specific optimal fusion strategies is a non-differentiable combinatorial optimization problem [24] where traditional gradient-based algorithms fail [41]. Thus, the crux lies in the structural inability of fixed architectures to autonomously optimize fusion strategies. Addressing this bottleneck necessitates shifting from manually designed architectures [40] to automated, task-specific architecture search [9]. While Evolutionary Algorithms (EAs) excel at such non-differentiable optimization problems [24], their deployment within MTL presents critical challenges. Existing MTL approaches often suffer from restricted shallow sharing, which lacks explicit channels for knowledge transfer [43, 19, 46, 8]. This hinders the cross-task exchange of advantageous architectural motifs among homologous targets, making it profoundly difficult to orchestrate complex interplays and avoid negative interference across diverse tasks. Therefore, constructing an efficient mechanism that simultaneously searches for optimal architectures and explicitly exploits biological homologies for cross-task “co-evolution” remains a crucial, unresolved hurdle for maximizing predictive robustness [44]. To address these challenges, this study develops the Multi-Task Genetic Algorithm with Multi-Granularity Encoding (MTGA-MGE), a unified framework specifically engineered to enhance protein-nucleotide binding site prediction. By synergizing adaptive architecture evolution with refined feature representation, MTGA-MGE systematically resolves the bottlenecks of rigid fusion and information redundancy. The main contributions of this work are as follows: (i) We design a Multi-Granularity Encoding (MGE) network to mitigate noise and redundancy inherent in high-dimensional embeddings. By synergizing multi-scale convolutions with self-attention mechanisms, MGE effectively distills discriminative local signals and captures global dependencies, thereby providing a highly refined representation for deciphering complex biological interactions. (i) We propose an NSGA-I-based Multi-Task Genetic Algorithm (MTGA) to bypass non-differentiable optimization constraints and overcome the limitations of static designs. This approach enables the adaptive evolution of task-specific feature fusion strategies, which improves generalization across diverse and complex binding scenarios. (i) We introduce an External-Neighborhood Mechanism (ENM) to catalyze targeted information exchange and facilitate efficient cross-task collaboration. By quantifying physicochemical homology to bridge task-specific channels, ENM strengthens co-evolution between tasks, ensuring robust predictive performance across diverse binding scenarios. I Related Work I-A Computational Methods for Protein-Nucleotide Binding Site Prediction The evolution of protein-nucleotide binding site prediction has transitioned from feature-engineered machine learning to end-to-end deep learning paradigms. Pioneering efforts focused on manual feature construction. For instance, EC-RUS [12] integrated evolutionary profiles with physicochemical properties via Random Forest classifiers, effectively enhancing representation diversity beyond simple sequence information. To ameliorate the pervasive class imbalance, BGSVM-NUC [49] employed Granular Support Vector Machines coupled with granular under-sampling techniques, effectively reducing model bias toward non-binding residues. With the advent of deep learning, research expanded into nucleotide-specific and nucleotide-general models. Nucleotide-specific models, such as DeepGTP [17] and DeepATPseq [14], leverage Convolutional Neural Networks (CNNs) to automatically extract deep local features, achieving superior precision for designated GTP and ATP targets, respectively. To broaden applicability, nucleotide-general models have emerged. SXGBsite [48] combines Extreme Gradient Boosting with SMOTE over-sampling to consolidate PSSM and solvent accessibility features, ensuring efficient and robust prediction. GHKNN [11] introduces a Graph Hashing-based K-Nearest Neighbors algorithm, which explicitly encodes topological dependencies between residues via graph structures to enhance spatial representation. NucGMTL [40] utilizes PLMs within a Grouped MTL framework, successfully extracting shared commonalities across diverse nucleotides to improve generalization. Figure 2: Framework of MTGA-MGE. The framework operates through three key components: I. Multi-Granularity Encoding, I. Multi-Task Genetic Algorithm, I. External-Neighborhood Mechanism. Finally, IV. Test Process illustrates the model inference phase. Nucleotide-specific models suffer from task isolation, precluding cross-task knowledge transfer and hindering performance on data-scarce ligands [22]. Conversely, although nucleotide-general models incorporate MTL, they predominantly rely on fixed, manually designed architectures. This static paradigm fails to efficiently navigate the vast discrete combinatorial space to identify optimal, task-specific fusion strategies [25]. These limitations underscore the urgent need to shift from manual architectural engineering to automated architecture search, which motivates our work. I-B Evolutionary Algorithms EAs, particularly Genetic Algorithms (GAs), have demonstrated robust global optimization capabilities across diverse complex systems, maintaining their vanguard status in modern deep learning optimization. Recently, the efficacy of EAs within bioinformatics has gained significant traction. For instance, Qian et al. [29] utilized evolutionary strategies to optimize deep fusion architectures for protein secondary structure prediction, demonstrating the superiority of automated architecture search over manual design. However, extrapolating these methodologies to protein-nucleotide binding site prediction presents unique challenges. Unlike the single-task nature of protein secondary structure prediction, nucleotide binding prediction necessitates MTL across diverse ligand types, thereby inducing an exponential expansion of the architecture search space [18]. To navigate such high-dimensional and discrete combinatorial landscapes, NSGA-I [15, 50, 47, 36] employs a reference-point-based selection mechanism. This mechanism ensures a uniform distribution of solutions along the Pareto front regardless of the search space’s continuity. Consequently, this mathematical property renders NSGA-I highly amenable to the discrete characteristics of the feature fusion problem addressed in this study [45]. I Methods This section provides an overview of the MTGA-MGE framework, followed by a detailed explanation of its three key components: the Multi-Granularity Encoding network, the Multi-Task Genetic Algorithm, and the External-Neighborhood Mechanism. The section concludes with the description of the test process for final inference. I-A Framework of MTGA-MGE The MTGA-MGE framework (Fig. 2) comprises three core components: Multi-Granularity Encoding (MGE), Multi-Task Genetic Algorithm (MTGA), and External-Neighborhood Mechanism (ENM). The MGE module constructs robust feature representations to overcome the challenges of high dimensionality and noise in features. It integrates multi-scale convolution with scale-based self-attention and employs a multi-task context injection mechanism. By doing so, it effectively distills discriminative residue-level patterns from redundant embeddings into a comprehensive candidate pool, enabling the model to more precisely pinpoint binding anchors within complex backgrounds. The MTGA module automates the architecture search process by formulating feature fusion as a combinatorial optimization problem. Utilizing the NSGA-I framework, it executes a series of evolutionary operations—such as crossover and mutation—to iteratively optimize the population of fusion strategies. Consequently, it adaptively navigates a complex search space to identify task-specific configurations defined by feature subsets, operators, and weights. This approach successfully addresses the inherent rigidity of manual designs, thereby effectively improving the overall generalization ability of the model. The ENM module introduces a collaborative knowledge interaction mechanism to enhance search efficiency and population diversity. It quantifies structural similarity via Gray Relational Analysis (GRA) to construct a dynamic reservoir of elite architectures from related tasks. Consequently, it explicitly guides the mutual exchange of valid structural blueprints, enabling efficient cross-task co-evolution. I-B MGE Method The MGE method comprises four main components: dual-view protein representation, multi-granularity feature extraction, loss function for class imbalance, and multi-task context injection. 1) Dual-view Protein Representation To construct a comprehensive feature space, we first employ ProtT5 [13] to capture global contextual information. This yields a robust semantic embedding matrix denoted as (t)∈ℝL×dtH^(t) ^L× d_t (dt=1024d_t=1024). Complementarily, we utilize ESM2 [21] to extract features rich in evolutionary signals. This model generates a high-dimensional structural representation denoted as (e)∈ℝL×deH^(e) ^L× d_e (de=1280d_e=1280). To synthesize these complementary perspectives, we concatenate the two embeddings along the channel dimension to form a unified dual-view representation ∈ℝL×dinH ^L× d_in, where din=dt+de=2304d_in=d_t+d_e=2304. This integration effectively synergizes the global semantic information from ProtT5 with the evolutionary structural patterns from ESM2, establishing a solid foundation for the subsequent MGE. 2) Multi-Granularity Feature Extraction To construct a diverse candidate pool and distill task-relevant patterns, we first define a dual-task context injection strategy to form the input tensor. Let T denote the set of 15 nucleotide-binding tasks, specifically =ADP, AMP, ATP, …, UTPT=\ADP, AMP, ATP, …, UTP\. Let τ∈ℝL×dinH_τ ^L× d_in (τ∈τ ) and τ′∈ℝL′×dinH_τ ^L × d_in (τ′∈τ (where τ≠τ′τ≠τ )) denote the unified dual-view feature matrices for the primary task τ and an auxiliary task τ′τ , respectively. We distinguish between single-task inputs (=τ∈ℝL×dinX=H_τ ^L× d_in) and dual-task inputs constructed by concatenating feature sets along the sequence dimension: =[τ∥τ′]∈ℝ(L+L′)×dinX=[H_τ _τ ] ^(L+L )× d_in. This high-dimensional input tensor X is then processed through the MGE module to extract multi-granularity features. The tensor is routed through three parallel branches, each characterized by a distinct kernel size k∈1,3,7k∈\1,3,7\ to capture local patterns at varying receptive fields. Unlike standard architectures that fuse multi-scale features prior to context modeling, our design integrates a 1D convolutional layer (dcnn=128d_cnn=128) followed directly by a specific Multi-Head Self-Attention (MHSA) module within each branch. Here, the MHSA employs the standard Transformer attention mechanism to capture long-range contextual dependencies across the sequence based on the extracted local representations. Reinforced by residual connections, the scale-specific refined features k∈ℝL×128Z_k ^L× 128 for each branch are computed as follows: k=MHSA(Conv1Dk())+Conv1Dk()Z_k=MHSA(Conv1D_k(X))+Conv1D_k(X) (1) where k∈1,3,7k∈\1,3,7\. This structure allows the model to capture long-range dependencies specifically tailored to different local granularities. Finally, the refined representations from all three branches are concatenated along the channel dimension and fused via a linear projection to produce the final deep representation ()∈ℝL×128E(X) ^L× 128: ()=Linear(1⊕3⊕7)E(X)=Linear(Z_1 _3 _7) (2) where ⊕ denotes the concatenation operation. This compact output serves as the robust feature basis for the subsequent evolutionary optimization stage. 3) Loss Function for Class Imbalance Given the extreme class imbalance between binding and non-binding residues, standard cross-entropy loss tends to bias the model toward the majority negative class. We therefore adopt the Focal Loss [20] to dynamically down-weight easy negatives and focus training on hard examples. For a residue with predicted probability p∈[0,1]p∈[0,1] and ground truth label y∈0,1y∈\0,1\, the loss is defined as: ℒfocal=−αy(1−p)γlog(p)−(1−α)(1−y)pγlog(1−p)L_focal=-α\,y(1-p)^γ (p)-(1-α)(1-y)p^γ (1-p) (3) where hyperparameters α and γ dynamically down-weight easy negatives and focus training on hard examples, effectively improving sensitivity for rare binding residues. 4) Dual-task Context Injection Single-task learning often ignores shared structural and evolutionary patterns across different nucleotide-binding tasks. To address this, we introduce a dual-task context injection strategy. By incorporating an auxiliary task, the encoder synthesizes complementary information, allowing the deep representation ()E(X) (defined in Eq. 2) to implicitly capture broader structural motifs. This constructs a diverse and information-rich candidate pool for the subsequent evolutionary search. Algorithm 1 Framework of MTGA 0: Data tr,valD_tr,D_val; Population size N; Max generations GmaxG_max. 0: Pareto optimal set for each task. 1: for each task t∈t do 2: Initialize PtP_t and evaluate fitness F()F(x); 3: end for 4: Initialize NSGA-I reference points Z; 5: for g=1g=1 to GmaxG_max do 6: ←N← Construct External-Neighborhoods via Algorithm 3; 7: for each task t∈t do 8: Retrieve tN_t from N; 9: Qt←∅Q_t← ; 10: while |Qt|<N|Q_t|<N do 11: c←c← Generate Offspring via Algorithm 2; 12: Qt←Qt∪cQ_t← Q_t∪\c\; 13: end while 14: Apply Batch DE on weights of QtQ_t; 15: Evaluate fitness F(Qt)F(Q_t); 16: Rt←Pt∪QtR_t← P_t∪ Q_t; 17: Perform Non-dominated Sort on RtR_t to get fronts; 18: Normalize objectives and associate with Z; 19: Pt←P_t← Select next generation from RtR_t via Niche-Preservation; 20: end for 21: end for 22: return Non-dominated individuals from final populations; For a target task τ∈τ , we generate a comprehensive feature pool containing one single-task representation and 28 dual-task representations by pairing τ with every other nucleotide task τ′∈τ (where τ≠τ′τ≠τ ). We explicitly employ a bidirectional pairing scheme—constructing both (τ,τ′)(τ,τ ) and (τ′,τ)(τ ,τ)—to capture asymmetric inter-task correlations, in each pair, only the first element (the primary task) serves as the supervision target to compute the training loss, whereas the second element (the auxiliary task) merely provides contextual features and does not participate in the loss calculation. Formally, the candidate input set poolX_pool is defined as: pool=τ∪⋃τ′∈∖τ[τ∥τ′],[τ′∥τ]X_pool=\H_τ\∪ _τ \τ\ \[H_τ _τ ],[H_τ _τ] \ (4) The resulting set of 29 candidate deep representations is thus given by: τ=()∣∈poolS_τ=\E(X) _pool\ (5) I-C MTGA Method To identify optimal task-specific feature fusion strategies, we formulate the learning problem for each nucleotide task as a bi-objective optimization. This problem inherently represents a mixed-integer programming challenge, involving both discrete combinatorial optimization (feature subset selection and operator arrangement) and continuous parameter optimization (fusion weight fine-tuning). 1) Bi-objective Formulation Algorithm 2 Offspring generation strategy 0: Current population PtP_t; External-Neighborhood tN_t; Transfer prob. PextP_ext; Crossover rate PcP_c; Mutation rate PmP_m; Specific probs Pstruct,Pop,Pweight\P_struct,P_op,P_weight\. 0: A generated offspring c. 1: Select first parent p1p_1 from PtP_t via Tournament Selection; 2: if rand()<Pextrand()<P_ext and p1≠∅N_p_1≠ then 3: Select second parent p2p_2 from p1N_p_1; 4: else 5: Select second parent p2p_2 from PtP_t; 6: end if 7: c←p1c← p_1.copy(); 8: if rand()<Pcrand()<P_c then 9: c←c← SinglePointCrossover(p1,p2)(p_1,p_2); 10: end if 11: if rand()<Pmrand()<P_m then 12: if rand()<Pstructrand()<P_struct then 13: Apply Structural Mutation; 14: end if 15: if rand()<Poprand()<P_op then 16: Apply Operator Mutation; 17: end if 18: if rand()<Pweightrand()<P_weight then 19: Apply Weight Mutation (using p1N_p_1); 20: end if 21: end if 22: return c; To efficiently evaluate fusion strategies, we employ a proxy mechanism where a logistic regression head is trained on standardized features for each individual x. This classifier utilizes the Focal Loss (defined in Eq. 3) to mitigate class imbalance. We formulate the optimization as maximizing the validation Area Under the Precision-Recall Curve (AUPRC, defined in Eq. 11) while minimizing the False Positive Rate (FPR). The FPR is defined as: FPR=FPFP+TNFPR= FPFP+TN (6) where FPFP and TNTN respectively represent the number of false positives and true negatives. Therefore, the bi-objective optimization problem is formally defined as: max f1()=AUPRC f_1(x)=AUPRC (7) min f2()=FPR f_2(x)=FPR The NSGA-I framework is then employed to optimize these conflicting objectives and maintain diversity along the Pareto front. 2) Semantically Aligned Feature Encoding Before detailing the individual representation and to lay the groundwork for valid cross-task transfer within the upcoming ENM, a unified representation space is strictly required. We propose a Semantically Aligned Encoding Scheme anchored on the fixed lexicographical order of the 15 nucleotide types (N0,…,N14N_0,…,N_14). Unlike relative indexing, the feature search space (indices 0−280-28, derived from Eq. 5) is organized by global biological identity. For any target task τ∈τ , the feature pool is structured as follows: Algorithm 3 Framework of ENM 0: Populations Pt\P_t\ and fitness for all tasks t∈t ; Neighborhood size K; Threshold ρ. 0: External-Neighborhood sets =tt∈N=\N_t\_t . 1: Initialize global elite pool ℰext←∅E_ext← ; 2: for each task t∈t do 3: Select elite individuals EtE_t from PtP_t based on primary objective; 4: ℰext←ℰext∪EtE_ext _ext∪ E_t; 5: end for 6: for each task t∈t do 7: t←∅N_t← ; 8: for each individual p∈Ptp∈ P_t do 9: Initialize similarity list Lp←∅L_p← ; 10: for each elite e∈ℰexte _ext from tasks t′≠t ≠ t do 11: Calculate Gray Relation Grade γ(p,e)γ(p,e); 12: Add (e,γ(p,e))(e,γ(p,e)) to LpL_p; 13: end for 14: Sort LpL_p in descending order of γ; 15: Select top-K elites from LpL_p to form pN_p; 16: end for 17: Add tN_t to N; 18: end for 19: return N; Indices 0–14 (Target as Primary): τ serves as the primary task, fused with each nucleotide k∈0−14k∈\0-14\ in the fixed set acting as auxiliary context. Notably, the index k where Nk=τN_k=τ denotes the single-task embedding (Self-Fusion), while others represent specific dual-task embeddings. Indices 15–28 (Target as Auxiliary): This range implements semantic role reversal, where τ serves as the auxiliary context supporting the other 14 nucleotides (sequentially ordered, excluding τ). This alignment enforces strict topological consistency: Index 0 invariably encodes the interaction with the first nucleotide (e.g., ADP), independent of the target task’s identity. This invariance ensures that beneficial interaction patterns identified by the evolutionary algorithm—such as leveraging ADP context—are semantically preserved during architectural migration. 3) Individual Representation and Feature Fusion Each individual employs a Feature-Operator-Weight (f-o-w) encoding comprising a subset of feature indices, a sequence of operators, and continuous weights. Let (0)()E^(0)(X) denote accumulated representation and (j)()∈τE^(j)(X) _τ (Eq. 2) the next selected feature matrix. The recursive fusion step is defined as: (0)()←wc⋅(0)()⊛wf⋅(j)()E^(0)(X)← w_c·E^(0)(X)\ \ w_f·E^(j)(X) (8) where ⊛∈add,mul,max,min,diff,avg ∈\ add, mul, max, min, diff, avg\. Sequential application of this operator over the selected subset yields the final fused matrix (fused)()E^(fused)(X). 4) Optimization Process The optimization framework (outlined in Algorithm 1) commences with random population initialization. In each generation, we first construct an external-neighborhood via GRA to facilitate valid cross-task knowledge transfer (detailed in Section I-D). Parents are then sampled from either the current task population or this constructed neighborhood, guided by the aligned encoding. To rigorously explore the search space, we apply a multi-level mutation strategy: Structural Mutation modifies the architecture topology (insertion/deletion/replacement); Operator Mutation updates interaction logic; and Weight Mutation perturbs continuous parameters. Crucially, to prevent the rejection of promising architectures due to suboptimal initialization, we incorporate a Batch Differential Evolution (Batch DE) mechanism. This mechanism employs a probabilistic DE/current-to-best/1 strategy. This lifetime learning step fine-tunes the continuous weights of offspring prior to fitness evaluation, ensuring architectures are assessed based on their potential optimal parameterization. Finally, the population is updated via NSGA-I environmental selection. I-D ENM Method Motivated by the shared physicochemical mechanisms (e.g., hydrogen bonding, electrostatics, and stacking) of structurally related ligands like ATP and GTP [43], we posit an architectural inductive bias: tasks governing biologically similar binding mechanisms imply similar optimal pathways for feature extraction. Consequently, we design the ENM that uses the Gray Relational Grade (GRG) [16] to explicitly measure strategy similarity. By constructing a reservoir of high-quality architectural motifs from related tasks, this mechanism allows the algorithm to transfer elite “structural blueprints” rather than searching from scratch, thereby facilitating collaborative exploration of the combinatorial search space. Inspired by MTEA/D-DN [37], our approach employs a unified feature–operator–weight (f-o-w) encoding (see Section 3) to enable direct similarity evaluation across tasks. We quantify individual similarity via GRG and construct the external-neighborhood for each task by selecting the most similar elite individuals from other tasks. 1) GRG Calculation Given individual vectors =[x1,…,xn]x=[x_1,…,x_n] and =[y1,…,yn]y=[y_1,…,y_n], we first compute the absolute deviation Δi=|yi−xi| _i=|y_i-x_i| for each dimension. Let Δmin _ and Δmax _ denote the minimum and maximum deviations across all dimensions. The Gray Relational Coefficient (GRC) for the i-th dimension is defined as: ζi=Δmin+ρΔmaxΔi+ρΔmax _i= _ +ρ _ _i+ρ _ (9) where the distinguishing coefficient ρ is set to 0.250.25. The GRG is then calculated as the mean GRC: γ(,)=1n∑i=1nζiγ(y,x)= 1n _i=1^n _i (10) 2) Elite Selection and ENM Process To mitigate the computational cost of exhaustive GRG calculation, we restrict the candidate pool to the elite subset of each task, defined as the set of top-performing individuals based on the primary objective (i.e., validation AUPRC). For a given individual, its external-neighborhood is formed by selecting the top-K most similar elites from other tasks based on GRG ranking (see Algorithm 3). TABLE I: Detailed statistics of the Nuc-1982 dataset distribution. Type Training and Validation Set Test Set (Pos, Neg)a Ratiob (Pos, Neg)a Ratiob ATP (4688, 183273) 2.55% (1403, 84267) 1.66% ADP (5145, 210355) 2.44% (1131, 68974) 1.64% AMP (2505, 100429) 2.49% (240, 12958) 1.85% GDP (1523, 56859) 2.67% (347, 16795) 2.06% GTP (1204, 50405) 2.38% (462, 19723) 2.34% TMP (224, 7652) 2.93% (31, 762) 4.07% CTP (484, 20888) 2.32% (152, 8706) 1.74% CMP (315, 8273) 3.81% (42, 3130) 1.34% UTP (444, 19712) 2.25% (56, 5801) 0.96% UMP (103, 2398) 4.29% (38, 1745) 2.18% UDP (974, 37356) 2.61% (230, 11954) 1.92% IMP (209, 6826) 3.06% (64, 2143) 2.99% GMP (178, 7060) 2.52% (9, 316) 2.85% CDP (311, 10922) 2.85% (58, 2559) 2.27% TTP (347, 14685) 2.36% (104, 7340) 1.42% aNumber of positive (Pos) and negative (Neg) samples. bRatio of positive samples. During offspring generation, the external-neighborhood serves as a reservoir of diverse genetic material. By allowing interactions between the current task population and its external-neighborhood during crossover and weight mutation operations—for instance, pairing a parent from the current task with a high-quality neighbor from a related task—this strategy introduces structurally valid yet diverse fusion patterns, thereby enhancing the algorithm’s global search capability and stability. I-E Test Process Upon conclusion of the evolutionary algorithms, the optimal fusion strategy is identified from the final Pareto front based on validation metrics and instantiated for inference. For a query protein sequence, the inference process commences by generating a candidate feature pool via the pretrained PLMs and the trained multi-granularity encoder. Guided by the decoded optimal strategy, specific feature subsets are selected, standardized, and sequentially fused according to the learned operators and weights. The resulting representation is mapped to logit scores via a pre-optimized coefficient matrix derived from the classification head, followed by a sigmoid activation to yield the final residue-level binding probabilities. IV EXPERIMENT DESIGN IV-A Datasets To rigorously evaluate model robustness and generalization under realistic scenarios, we employ the widely recognized and comprehensive Nuc-1982 benchmark dataset. This dataset encompasses 15 nucleotide-binding tasks, comprising five high-resource nucleotides (e.g., ATP, GTP) and ten low-resource nucleotides. TABLE I: Hyperparameters used in MTGA-MGE Hyperparameter Setting Stage I: Feature Extraction Model Architecture Input Dimension 2304 (ProtT5 + ESM2) CNN Kernel Sizes 1, 3, 7 CNN Channel Dimensions 512 → 256 → 128 Transformer Heads 8 Transformer Hidden Dim 128 Dropout Rate 0.5 Optimization & Setup Optimizer AdamW Learning Rate 1×10−41× 10^-4 Weight Decay 1×10−41× 10^-4 Batch Size 16 Maximum Epochs 30 Early Stopping Patience 5 Validation Set Ratio 25% Loss Function Type Focal Loss Focusing Parameter (γ) 1.5 Class Balancing Weights (α) Negative: 0.15, Positive: 0.85 Stage I: Evolutionary Algorithms Evolutionary Strategy Optimizer NSGA-I Population Size 50 Number of Generations 40 Crossover Probability 0.9 Mutation Probability 0.6 Maximum Feature Length 25 Transfer Neighborhood Size (K) 10% of Population Gray Relation Coefficient (ρ) 0.25 Evaluation Head (Logistic) Maximum Iterations 300 Ridge Regularization (λ) 0.5 Data Partitioning and Preprocessing: We adopted a strict temporal split strategy to simulate predictions on future data. The training and validation sets were derived from proteins released prior to October 26, 2023. To mitigate bias arising from homologous proteins, sequences within each nucleotide category were redundancy-reduced at a 30% sequence identity threshold. For the three data-heavy tasks (ADP, AMP, ATP), we implemented a 25% stratified subsampling to balance computational efficiency with representativeness, while retaining all available data for the remaining tasks. Preliminary experiments showed that this subsampling ratio provided sufficient representativeness for convergence without compromising performance stability. Ultimately, 25% of the processed data was held out as the validation set. The independent test set consists of proteins released on or after October 26, 2023. This time-based partitioning ensures zero data leakage during evaluation, thereby objectively reflecting the model’s generalization boundaries. Distributional Challenge: The benchmark is characterized by extreme Class Imbalance (Table I). Positive binding residues account for only ∼ 2.77% and 2.09% of the training and independent test sets, respectively. This sparse distribution poses a severe challenge for the model to capture faint binding signals while maintaining a low false positive rate. TABLE I: Performance Comparison with State-of-the-Art Methods on Five High-resource Nucleotide Tasks Nucleotide Method MCC AUPRC Sen (%) Pre (%) Acc (%) Spe (%) ATP SXGBsite 0.343 0.377 52.49 24.62 96.16 96.98 EC-RUS 0.423 0.311 48.77 38.78 97.64 98.55 BGSVM-NUC 0.357 0.350 45.94 30.01 97.03 97.99 GHKNN 0.120 0.223 46.09 5.67 84.89 85.62 DeepATPseq 0.531 0.506 49.11 58.94 98.61 99.43 NucGMTL 0.570 0.552 55.38 60.10 98.66 99.38 MTGA-MGE 0.667 0.716 69.85 64.86 98.89 99.37 ADP SXGBsite 0.387 0.425 57.19 28.03 96.87 97.53 EC-RUS 0.484 0.390 55.85 43.55 98.08 98.78 BGSVM-NUC 0.405 0.390 50.22 34.65 97.61 98.41 GHKNN 0.162 0.314 51.30 7.37 88.55 89.17 NucGMTL 0.604 0.599 59.84 62.32 98.76 99.40 MTGA-MGE 0.684 0.674 50.13 94.19 99.15 99.95 AMP SXGBsite 0.279 0.269 50.00 17.99 94.95 95.78 EC-RUS 0.392 0.259 37.50 43.06 97.96 99.08 BGSVM-NUC 0.252 0.224 29.58 24.15 97.03 98.27 GHKNN 0.111 0.194 33.33 6.44 89.98 91.02 NucGMTL 0.555 0.572 50.58 62.62 98.55 99.44 MTGA-MGE 0.619 0.648 67.50 58.06 98.52 99.10 GDP SXGBsite 0.488 0.466 50.15 49.85 97.57 98.74 EC-RUS 0.506 0.499 36.17 73.01 98.13 99.67 BGSVM-NUC 0.494 0.447 39.82 63.59 98.00 99.43 GHKNN 0.323 0.423 50.46 23.71 94.87 95.97 NucGMTL 0.661 0.654 61.56 72.24 98.74 99.51 MTGA-MGE 0.695 0.754 61.38 79.78 98.90 99.68 GTP SXGBsite 0.357 0.362 40.04 34.64 96.90 98.23 EC-RUS 0.399 0.384 41.34 41.16 97.30 98.62 BGSVM-NUC 0.415 0.352 32.25 56.02 97.87 99.41 GHKNN 0.270 0.349 40.04 21.39 95.26 96.55 NucGMTL 0.572 0.594 50.48 66.72 98.29 99.41 MTGA-MGE 0.649 0.675 65.58 65.87 98.43 99.20 Figure 3: Radar plots comparing the performance of NucMTL (red), NucGMTL (blue), and MTGA-MGE (green) on ten low-resource nucleotide tasks. Panels (a)–(c) display the MCC, while panels (d)–(f) present the AUPRC. In each plot, the axes represent specific nucleotide tasks, with a larger enclosed area indicating superior predictive performance. (a) 7pt7_8 (b) 7vuk_B (c) 7wak_A (d) 7zvs_A (e) 7s5e_A (f) 7r30_B (g) 7w1m_A (h) 8ak3_A Figure 4: Visualization of protein-ADP binding site predictions. The protein structures are shown in gray cartoon representation. Red: True Positive (TP); Blue: False Negative (FN); Yellow: False Positive (FP). IV-B Parameter Settings All experiments were conducted using the PyTorch framework on a single NVIDIA GeForce RTX 4090 GPU. In the feature extraction stage, the encoder is optimized using AdamW, employing Focal Loss to address the severe class imbalance. For the evolutionary stage, we utilize the NSGA-I algorithm to balance feature dimensionality and predictive accuracy. Detailed configurations for both model architecture and optimization strategies are comprehensively listed in Table I. IV-C Evaluation The model’s performance is assessed using six key metrics: AUPRC, Matthews Correlation Coefficient (MCC), Sensitivity (Sen), Precision (Pre), Specificity (Spe), and Accuracy (Acc). AUPRC and MCC: AUPRC and MCC are prioritized as the primary metrics, given that the extreme class imbalance (average positive ratio <3%<3\%) renders standard accuracy unreliable for residue-level prediction. AUPRC quantifies the trade-off between precision and recall across decision thresholds: AUPRC=∫01P(R)RAUPRC= _0^1P(R)\,dR (11) where P(R)P(R) denotes precision as a function of recall R. Additionally, MCC provides a robust quality measure even for highly unbalanced classes by utilizing all confusion matrix entries: MCC=TP⋅TN−FP⋅FN(TP+FP)(TP+FN)(TN+FP)(TN+FN)MCC= TP· TN-FP· FN (TP+FP)(TP+FN)(TN+FP)(TN+FN) (12) where TPTP, TNTN, FPFP, and FNFN represent true positives, true negatives, false positives, and false negatives, respectively. Supplementary Metrics: To ensure comprehensive evaluation, we also report Sensitivity, Precision, Specificity, and Accuracy, calculated as follows: Sen=TPTP+FNSen= TPTP+FN (13) Pre=TPTP+FPPre= TPTP+FP (14) Spe=TNTN+FPSpe= TNTN+FP (15) Acc=TP+TNTP+TN+FP+FNAcc= TP+TNTP+TN+FP+FN (16) V Experimental Results V-A Comparison with State-of-the-Art Methods We rigorously benchmarked MTGA-MGE against six representative nucleotide-general models and several nucleotide-specific predictors across 15 tasks in the Nuc-1982 suite. The evaluation covered both high-resource and low-resource nucleotides, with detailed results presented in Table I and Fig. 3. Additionally, the spatial distribution of predictions is qualitatively illustrated in Fig. 4, where the predicted residues (red) accurately delineate the physical binding pockets. Figure 5: Overall performances of MGE. (1) Performance comparison (MCC and AUPRC) of different convolutional encoder architectures across five high-resource nucleotide tasks. (2) t-SNE visualization comparing the residue representations learned (a) using raw PLM features alone and (b) using MGE. (3) t-SNE visualization comparing residue representations learned under different input configurations: (a) single-task GDP input and (b) joint GDP–ADP input. In all t-SNE plots, data points are colored by prediction outcome: true positive (TP, purple), false negative (FN, red), false positive (FP, blue), and true negative (TN, gray). Performance on High-resource Nucleotide Tasks: MTGA-MGE achieved the highest AUPRC and MCC scores across all five high-resource nucleotide tasks, significantly outperforming state-of-the-art competitors. As shown in Table I, our model surpassed the nearest baseline by an average margin of 0.099 in AUPRC and 0.070 in MCC. Task-specific analysis reveals that for AMP, the high AUPRC despite lower precision indicates an effective ranking of true binding residues; conversely, for ADP, the combination of high precision and robust sensitivity suggests the establishment of strict and accurate decision boundaries within highly skewed distributions. Performance on Low-resource Nucleotide Tasks: Despite extreme data scarcity, MTGA-MGE demonstrated competitive adaptability when compared to both the traditional multi-task learning baseline (NucMTL) and the state-of-the-art NucGMTL (Fig. 3). Among the ten low-resource nucleotide tasks, the model secured performance gains on six, avoiding performance collapse under low-resource conditions. Notably, even in the extremely data-scarce UTP and CTP tasks, MTGA-MGE maintained competitive AUPRC benchmarks. V-B Effectiveness of MGE To validate the architectural advantage of the MGE, we conducted a controlled ablation study treating the encoder as the sole variable. Under a unified training protocol, we benchmarked our MGE against three representative variants: a single-scale baseline, a serial stacked configuration, and a dual-branch parallel architecture. This setup disentangles the contribution of different receptive field modeling strategies to representation quality. As illustrated in Fig. 5(1), quantitative analysis confirms the consistent superiority of the MGE across all five high-resource nucleotide tasks. The single-scale baseline exhibited the poorest performance, verifying that limited receptive fields are inadequate for capturing complex binding cues under extreme class imbalance. While serial and simple dual-branch configurations offer marginal improvements, MGE achieves the highest MCC and AUPRC scores. This demonstrates that the parallel integration of complementary semantics from short-, medium-, and long-range contexts is optimal for robust feature extraction. Figure 6: Overall performance of MTGA optimization and MTGA-MGE inference efficiency. (1) Performance comparison between Naive Mean Fusion (MGE-Avg) and GA (MTGA-MGE) on five high-resource nucleotide tasks. (2) Temporal evolution of interaction intensity for the GDP task across evolutionary generations. (3) Evolution of the Pareto front for the GDP task. The left panel shows the initial population with random scattering, and the right panel illustrates the final population converging toward a non-dominated front. (4) Comparison of inference efficiency between NucGMTL and MTGA-MGE on the ADP task across different batch sizes. Visualization Analysis: To intuitively elucidate the encoder’s role, we employed t-SNE to visualize the latent representations (Fig. 5(2)). Compared to raw PLM embeddings, the refined features exhibit markedly enhanced class separability, characterized by distinct clustering of positive and negative samples. This qualitatively corroborates the encoder’s efficacy in filtering high-dimensional background noise. Notably, the error distribution reveals a specific “halo effect”: false positives are not randomly scattered but are spatially clustered in the immediate vicinity of true binding sites. This phenomenon reflects an intrinsic trade-off in convolutional modeling: while multi-scale aggregation amplifies weak binding signals (thereby boosting recall), it inevitably blurs local boundaries via spatial smoothing. These “near-misses” indicate that the model successfully localizes binding pockets despite imprecise boundary delineation, suggesting future avenues for fine-grained boundary refinement. V-C Effectiveness of Dual-task Context Injection To validate the benefits of the proposed multi-task context injection, we compared feature representations derived from single-task GDP versus joint GDP–ADP inputs. As visualized in Fig. 5(3), the joint-input configuration qualitatively exhibits markedly clearer class separation, an observation rigorously corroborated by quantitative metrics computed on the original high-dimensional embeddings. The average inter-class distance increased from 10.8610.86 to 13.8113.81, while the Davies–Bouldin (DB) score—where a lower value indicates superior discriminability based on the ratio of intra-class compactness to inter-class separation—improved from 0.9480.948 to 0.7210.721. These results confirm that joint-task training effectively leverages cross-task semantic correlations to enhance feature discriminability. This performance gain can be attributed to the underlying biological homology. Since homologous ligands (e.g., GDP and ADP) typically target pockets with similar physicochemical properties, we hypothesize that the auxiliary task provides a critical “structural reference” during training. For residues exhibiting ambiguous signals in the target task, the shared auxiliary context likely facilitates the alignment of their features toward regions with well-defined binding patterns. This process appears to effectively mitigate feature divergence in single-task scenarios, transforming loose, task-specific distributions into a more coherent and collaborative representation. V-D Impact of GA To assess the contribution of GA, we benchmarked the full MTGA-MGE framework against a non-evolutionary baseline, MGE-Avg, where the adaptive genetic algorithm is replaced by a static arithmetic mean strategy across all 29 candidate feature sources. As visualized in Fig. 6(1), replacing evolutionary selection with naive aggregation induces severe performance degradation across all high-resource nucleotide tasks. This failure is particularly catastrophic for GDP and GTP; notably, MGE-Avg achieves a negligible MCC of 0.051 on GDP (compared to 0.695 for MTGA-MGE), essentially losing all discriminative capability. TABLE IV: Performance Comparison between With and Without the ENM on Five High-resource Nucleotides Nucleotide Optimization Strategy MCC AUPRC ADP w/o-neighborhood 0.596 0.611 neighborhood 0.593 0.630 AMP w/o-neighborhood 0.567 0.530 neighborhood 0.618 0.600 ATP w/o-neighborhood 0.653 0.699 neighborhood 0.677 0.715 GDP w/o-neighborhood 0.653 0.670 neighborhood 0.700 0.742 GTP w/o-neighborhood 0.586 0.603 neighborhood 0.637 0.665 This collapse highlights the inherent risks of static multi-task aggregation. By indiscriminately averaging the target task with 28 auxiliary contexts, the baseline treats all information sources as equally relevant. Consequently, the distinctive feature patterns required for identifying specific binding residues are likely compromised by irrelevant or conflicting information from biologically distinct tasks. In contrast, MTGA-MGE’s superiority stems from its capacity for autonomous selective fusion. By dynamically retaining only compatible feature subsets, the evolutionary framework effectively filters out interference while amplifying target signals. These results confirm that the adaptive selection provided by the evolutionary module, rather than mere information accumulation, is indispensable for resolving task heterogeneity and ensuring robust performance. V-E Impact of ENM To evaluate the contribution of ENM, we conducted a controlled ablation study on five high-resource nucleotide tasks. We contrast the baseline configuration lacking the external mechanism (denoted as w/o-neighborhood) against the enhanced strategy incorporating cross-task neighborhoods (denoted as neighborhood). In the w/o-neighborhood setting, each task evolves independently based solely on its own population; conversely, the neighborhood setting permits offspring to sample high-quality parents from the external-neighborhood, facilitating selective cross-task genetic variation. All other experimental configurations were kept identical to ensure rigorous comparability. As evidenced in Table IV, neighborhood consistently outperforms w/o-neighborhood in AUPRC and MCC metrics across four of the five tasks, with only marginal fluctuations observed for ADP. These results suggest that by exploiting complementary search trajectories from biologically related tasks, the neighborhood strategy effectively bolsters the model’s capacity to detect rare binding residues under conditions of extreme class imbalance. To elucidate the underlying dynamics of this gain, we monitored the temporal evolution of interaction intensity for the GDP task (Fig. 6(2)). The process exhibits distinct phasic characteristics: initial generations show high-frequency interactions with external tasks (ADP, ATP, GTP), indicating that the ENM actively leverages diversity to escape local optima during the early exploration phase. As evolution proceeds, reliance on cross-task transfer diminishes in favor of task-specific refinement. This transition confirms that the mechanism facilitates a natural shift from global collaborative exploration to local exploitation, autonomously regulating the dependency on information transfer as the population converges. TABLE V: Performance Comparison of Different Multi-Objective Optimization Strategies on Four Representative High-resource Nucleotide Tasks Nucleotide Method MCC AUPRC ADP SPEA2 0.610 0.603 NSGA-I 0.586 0.544 NSGA-I 0.684 0.674 AMP SPEA2 0.681 0.646 NSGA-I 0.659 0.643 NSGA-I 0.684 0.648 GDP SPEA2 0.645 0.585 NSGA-I 0.694 0.729 NSGA-I 0.700 0.742 GTP SPEA2 0.581 0.574 NSGA-I 0.574 0.565 NSGA-I 0.592 0.601 V-F Comparison of Different Multi-Objective Optimizers To justify the selection of NSGA-I as the core optimization engine, we conducted a controlled comparison against two established baselines: NSGA-I, which relies on crowding distance for diversity maintenance, and SPEA2, which utilizes a strength-based fitness and external archive mechanism. Quantitative results across four representative high-resource nucleotide tasks (ADP, AMP, GDP, GTP), summarized in Table V, demonstrate that NSGA-I consistently yields the highest MCC and AUPRC scores. This superiority is likely attributed to the unique reference-point-based mechanism of NSGA-I. Unlike the crowding distance estimation of NSGA-I or the nearest-neighbor density estimation of SPEA2, the reference-point strategy provides a more structured and uniform exploration of the solution space, making it particularly robust for navigating the complex discrete combinatorial landscape of feature fusion. Fig. 6(3) tracks the evolution of the solution space for the GDP task. The transition from an initially scattered distribution to a condensed, continuous frontier empirically validates the algorithm’s convergence capability. This progression from a diffuse cloud to a well-defined frontier demonstrates effective selection pressure, successfully driving the population toward non-dominated solutions that represent optimal trade-offs between sensitivity (AUPRC) and specificity (FPR) for downstream model selection. V-G Inference Efficiency Analysis To evaluate the practical utility of MTGA-MGE, we benchmarked its inference latency against the state-of-the-art NucGMTL model. The ADP-binding task was selected as the representative benchmark because its number of fused features lies in the median range among all evaluated nucleotides, providing a balanced perspective on computational complexity and model performance. Although MTGA-MGE involves an evolutionary search during the training phase, the final inference is optimized by pruning the computational graph to execute only the task-specific optimal paths. The empirical results, as illustrated in Fig. 6(4), demonstrate that MTGA-MGE maintains high efficiency across different batch sizes. Notably, for medium-scale inputs (10–50 sequences), MTGA-MGE even outperforms NucGMTL, completing 50 sequences in 18.36 seconds compared to 26.80 seconds. These results verify that our proposed multi-granularity encoding and adaptive architecture do not impose a significant computational burden during deployment, ensuring its viability for high-throughput protein-nucleotide binding site prediction. VI Discussion Based on comprehensive evaluations across the Nuc-1982 benchmark and various cross-ligand scenarios, we conclude that the MTGA-MGE framework offers distinct advantages in three key aspects: (1) Superior Predictive Robustness and Generalization. By transitioning from manual feature fusion to an adaptive optimization paradigm, the model effectively breaks the limitations inherent in fixed architectures. This flexibility ensures that the learned representations remain robust when faced with the complex structural characteristics of low-resource nucleotides, thereby effectively enhancing the model’s overall generalization capability. (2) Enhanced Representation Purity and Signal Sensitivity. This study achieves a significant leap in representational clarity by precisely distilling core binding signals from complex backgrounds. By refining the granularity of feature extraction, the model filters out high-dimensional noise that often obscures critical residues. This refinement heightens the model’s sensitivity to the subtle spatial arrangements of binding sites, enabling more precise localization of interaction anchors within intricate protein structures. (3) Homology-Driven Biological Synergy. The framework effectively bridges the gap between isolated tasks by explicitly exploiting physicochemical homologies. By transforming disjoint, task-specific distributions into structurally aligned representations, the model allows the shared biological context to calibrate latent features. This “co-evolutionary” design facilitates structural information complementarity across tasks, thereby stabilizing the predictive performance of various ligands and maximizing the overall robustness of the framework. VII Conclusion In this study, we addressed the challenges of task heterogeneity in protein-nucleotide binding site prediction by proposing MTGA-MGE, a unified framework that transitions from rigid, manual feature fusion to an autonomous, adaptive architecture search. Extensive evaluations demonstrate that MTGA-MGE establishes a new state-of-the-art on high-resource nucleotide tasks while maintaining competitiveness on low-resource ones. These results validate that modeling feature fusion as a discrete combinatorial optimization problem is a powerful approach to navigating the semantic space of high-dimensional multi-source representations. Despite these strengths, we acknowledge certain limitations that warrant further investigation. First, the evolutionary process incurs high computational overhead during training due to the iterative nature of the search. However, this cost is strictly confined to the offline stage; once the architecture is optimized, online inference remains exceptionally fast and viable for high-throughput screening. Second, the model remains sensitive to extreme data scarcity. Effective transfer requires minimal distributional overlap, and in regimes of extreme sparsity, stochastic noise may outweigh transfer benefits. Future work will address these constraints by integrating structural priors or few-shot learning techniques to elevate performance in low-resource scenarios, and we plan to explore surrogate-assisted optimization to accelerate the search process for broader ligand-binding domains. References [1] E. E. Alemasova and O. I. Lavrik (2017) At the interface of three nucleic acids: the role of RNA-binding proteins and poly (ADP-ribose) in DNA repair. Acta Naturae 9 (2), p. 4–16. Cited by: §I. [2] A. J. Berdis (2022) Nucleobase-modified nucleosides and nucleotides: applications in biochemistry, synthetic biology, and drug discovery. Front. Chem. 10, p. 1051525. External Links: Document Cited by: §I. [3] B. E. Blass (2015) Basic principles of drug discovery and development. Elsevier, Oxford, UK. Cited by: §I. [4] O. Boutureira and G. J. L. Bernardes (2015) Advances in chemical protein modification. Chem. Rev. 115 (5), p. 2174–2195. Cited by: §I. [5] N. Brandes, D. Ofer, et al. (2022) ProteinBERT: a universal deep-learning model of protein sequence and function. Bioinformatics 38 (8), p. 2102–2110. Cited by: §I. [6] N. S. Chandel (2021) Nucleotide metabolism. Cold Spring Harb. Perspect. Biol. 13 (7), p. a040592. External Links: Document Cited by: §I. [7] J. S. Chauhan, N. K. Mishra, et al. (2010) Prediction of GTP interacting residues, dipeptides and tripeptides in a protein from its evolutionary information. BMC Bioinform. 11 (1), p. 301. Cited by: §I. [8] J. Chen, Q. Xu, et al. (2024) Graph-based multi-task learning for protein-DNA binding site prediction. Nucleic Acids Res. 52 (4), p. e21. Cited by: §I. [9] Q. Danneels and M. Verbeke (2025) MTL-simnas: task similarity-driven neural architecture search for enhanced multi-task learning. In Int. Conf. Artif. Neural Netw., p. 272–284. Cited by: §I. [10] V. de Crécy-Lagard, R. Dias, et al. (2025) Limitations of current machine learning models in predicting enzymatic functions for uncharacterized proteins. G3 (Bethesda) 15 (10), p. jkaf169. Cited by: §I. [11] Y. s. Ding, C. Yang, et al. (2022) Identification of protein-nucleotide binding residues via graph regularized k-local hyperplane distance nearest neighbor model. Appl. Intell. 52 (6), p. 6598–6612. Cited by: §I, §I-A. [12] Y. Ding, J. Tang, et al. (2017) Identification of protein-ligand binding sites by sequence information and ensemble classifier. J. Chem. Inf. Model. 57 (12), p. 3149–3161. Cited by: §I, §I-A. [13] A. Elnaggar, M. Heinzinger, et al. (2021) ProtTrans: towards cracking the language of life’s code through self-supervised learning. IEEE Trans. Pattern Anal. Mach. Intell. 44, p. 7112–7127. Cited by: §I, §I-B. [14] J. Hu, L. Zheng, et al. (2021) Accurate prediction of protein-ATP binding residues using position-specific frequency matrix. Anal. Biochem. 626, p. 114241. Cited by: §I, §I-A. [15] M. Jiang, W. Wang, et al. (2023) Multi-objective evolutionary architecture search for protein-protein interaction prediction. Inf. Sci. 622, p. 1012–1028. Cited by: §I-B. [16] N. Kandula (2023) Gray relational analysis of tuberculosis drug interactions: a multi-parameter evaluation of treatment efficacy. J. Comput. Sci. Appl. Inf. Technol.. Cited by: §I-D. [17] Q. Le, Q. Ho, et al. (2019) Using two-dimensional convolutional neural networks for identifying GTP binding sites in Rab proteins. J. Bioinform. Comput. Biol. 17 (1), p. 1950005. Cited by: §I, §I-A. [18] H. Li and T. Liu (2026) Contribution-driven task design: multi-task optimization algorithm for large-scale constrained multi-objective problems. Computers 15 (1), p. 31. Cited by: §I-B. [19] S. Li, F. Guo, et al. (2022) DeepMapi: deep multi-task learning for predicting protein–nucleic acid binding residues. Brief. Bioinform. 23 (1), p. bbab502. Cited by: §I. [20] T. Lin, P. Goyal, et al. (2017) Focal loss for dense object detection. In Proc. IEEE Int. Conf. Comput. Vis. (ICCV), p. 2980–2988. Cited by: §I-B. [21] Z. Lin, H. Akin, et al. (2023) Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379 (6637), p. 1123–1130. Cited by: §I, §I-B. [22] J. Liu, Y. Ding, et al. (2026) DTAP: a unified graph transformer framework for joint prediction of drug–target affinity and docking pose. Brief. Bioinform. 27 (1), p. bbag069. Cited by: §I-A. [23] S. Liu, Y. Liang, et al. (2019) Loss-balanced task weighting to reduce negative transfer in multi-task learning. In Proc. AAAI Conf. Artif. Intell., Vol. 33, p. 9977–9978. Cited by: §I. [24] Y. Liu, Y. Sun, et al. (2023) A survey on evolutionary neural architecture search. IEEE Trans. Neural Netw. Learn. Syst. 34 (2), p. 550–570. Cited by: §I, §I. [25] K. Long, G. Xie, et al. (2025) Enhancing multimodal learning via hierarchical fusion architecture search with inconsistency mitigation. IEEE Trans. Image Process.. Cited by: §I-A. [26] N. NaderiAlizadeh and R. Singh (2025) Aggregating residue-level protein language model embeddings with optimal transport. Bioinform. Adv. 5 (1), p. vbaf060. Cited by: §I. [27] M. Pourmirzaei, S. Alqarghuli, et al. (2025) Zero-shot protein–ligand binding site prediction from protein sequence and smiles. bioRxiv. Cited by: §I. [28] M. Pourmirzaei, F. Esmaili, et al. (2025) Prot2token: a unified framework for protein modeling via next-token prediction. arXiv preprint arXiv:2505.20589. Cited by: §I. [29] Y. Qian, L. Su, et al. (2025) Enhanced protein secondary structure prediction through multi-view multi-feature evolutionary deep fusion method. IEEE Trans. Emerg. Top. Comput. Intell.. Cited by: §I-B. [30] C. Rhodes, S. Balaratnam, et al. (2024) Targeting rna-protein interactions with small molecules: promise and therapeutic potential. Med. Chem. Res. 33 (11), p. 2050–2065. Cited by: §I. [31] J. M. Sagendorf, H. M. Berman, et al. (2017) DNAproDB: an interactive tool for structural analysis of DNA–protein complexes. Nucleic Acids Res. 45 (W1), p. W89–W97. Cited by: §I. [32] A. Stützer, L. M. Welp, et al. (2020) Analysis of protein–DNA interactions in chromatin by UV induced cross-linking and mass spectrometry. Nat. Commun. 11 (1), p. 5250. Cited by: §I. [33] E. Valkov, T. Sharpe, et al. (2011) Targeting protein–protein interactions and fragment-based drug discovery. In Fragment-Based Drug Discovery and X-Ray Crystallography, p. 145–179. Cited by: §I. [34] J. van Eck, D. Gogishvili, et al. (2026) PLM-explain: divide and conquer the protein embedding space. Bioinformatics 42 (1), p. btaf631. Cited by: §I. [35] R. van Tienhoven and A. Zaldumbide (2025) The dark proteome: “not everything that counts can be counted”. Diabetes 74 (12), p. 2211–2213. Cited by: §I. [36] S. Wang, P. Cui, et al. (2026) Enhanced protein intrinsic disorder prediction through dual-view multiscale features and multi-objective evolutionary algorithm. Note: arXivarXiv preprint arXiv:2603.06292 External Links: Document, Link Cited by: §I-B. [37] X. Wang, Z. Dong, et al. (2023) Multiobjective multitask optimization-neighborhood as a bridge for knowledge transfer. IEEE Trans. Evol. Comput. 27 (1), p. 155–169. Cited by: §I-D. [38] Y. Wang, F. Boadu, et al. (2025) MPBind: a multitask protein binding site predictor using protein language models and equivariant gnns. Bioinformatics 41 (11), p. btaf589. Cited by: §I. [39] J. Wu, F. Ge, et al. (2025) Identification of protein-nucleotide binding residues with deep multi-task and multi-scale learning. IEEE J. Biomed. Health Inform.. Cited by: §I, §I. [40] J. Wu, Y. Liu, et al. (2025) Identifying protein-nucleotide binding residues via grouped multi-task learning and pre-trained protein language models. J. Chem. Inf. Model. 65 (2), p. 1040–1052. Cited by: §I, §I, §I, §I-A. [41] Y. Xue, K. Liu, et al. (2025) Dominant classifier-assisted hybrid evolutionary multi-objective neural architecture search.. Int. J. Neural Syst. 35 (10), p. 2550051–2550051. Cited by: §I. [42] J. Yan, S. Friedrich, et al. (2016) A comprehensive comparative review of sequence-based predictors of DNA- and RNA-binding residues. Brief. Bioinform. 17 (1), p. 88–105. Cited by: §I. [43] H. Zhang, Y. Li, et al. (2022) Deep learning for pan-specific protein-nucleotide binding site prediction. Bioinformatics 38 (7), p. 1841–1849. Cited by: §I, §I-D. [44] T. Zhang, D. Li, et al. (2024) Constrained multitasking optimization via co-evolution and domain adaptation. Swarm Evol. Comput. 87, p. 101570. Cited by: §I. [45] Y. Zhang et al. (2025) Discrete neural architecture search via reference-point-based evolutionary multi-objective optimization. IEEE Trans. Evol. Comput.. Cited by: §I-B. [46] Y. Zhang, S. Wang, et al. (2023) Cross-task knowledge distillation for protein-ligand interaction prediction. Bioinformatics 39 (5), p. btad214. Cited by: §I. [47] F. Zhao, K. Tang, et al. (2025) Adaptive reference point NSGA-I for large-scale discrete combinatorial optimization. Eur. J. Oper. Res. 320 (1), p. 45–62. Cited by: §I-B. [48] Z. Zhao, Y. Xu, et al. (2019) SXGBsite: prediction of protein-ligand binding sites using sequence information and extreme gradient boosting. Genes 10 (12), p. 965. Cited by: §I, §I-A. [49] Y. Zhu, J. Hu, et al. (2019) Boosting granular support vector machines for the accurate prediction of protein-nucleotide binding sites. Comb. Chem. High Throughput Screen. 22 (7), p. 455–469. Cited by: §I, §I-A. [50] Z. Zhu, M. Gong, et al. (2024) Evolutionary multi-objective optimization for high-dimensional feature fusion in bioinformatics. Evol. Comput. 32 (2), p. 189–210. Cited by: §I-B.