Paper deep dive
CAPSUL: A Comprehensive Human Protein Benchmark for Subcellular Localization
Yicheng Hu, Xinyu Lin, Shulin Li, Wenjie Wang, Fengbin Zhu, Fuli Feng
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 6:06:19 AM
Summary
The paper introduces CAPSUL, a comprehensive human protein benchmark for subcellular localization that integrates 3D structural information (AlphaFold2, FoldSeek) with fine-grained, expert-curated localization annotations from UniProt and HPA. It addresses the limitations of existing datasets like DeepLoc by providing 20 distinct subcellular compartments and experimental evidence levels, enabling the evaluation of both sequence-based and structure-based protein representation models.
Entities (6)
Relation Signals (3)
CAPSUL → contains → 20 subcellular compartments
confidence 95% · we further refine the subcellular area space by introducing 20 aggregated subcellular compartments
CAPSUL → integrates → 3D structural information
confidence 95% · It features a dataset that integrates diverse 3D structural representations with fine-grained subcellular localization annotations
AlphaFold2 → providesdatafor → CAPSUL
confidence 95% · we leverage AlphaFold2 to extract the Cartesian coordinates
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Subcellular localization is a crucial biological task for drug target identification and function annotation. Although it has been biologically realized that subcellular localization is closely associated with protein structure, no existing dataset offers comprehensive 3D structural information with detailed subcellular localization annotations, thus severely hindering the application of promising structure-based models on this task. To address this gap, we introduce a new benchmark called $\mathbf{CAPSUL}$, a $\mathbf{C}$omprehensive hum$\mathbf{A}$n $\mathbf{P}$rotein benchmark for $\mathbf{SU}$bcellular $\mathbf{L}$ocalization. It features a dataset that integrates diverse 3D structural representations with fine-grained subcellular localization annotations carefully curated by domain experts. We evaluate this benchmark using a variety of state-of-the-art sequence-based and structure-based models, showcasing the importance of involving structural features in this task. Furthermore, we explore reweighting and single-label classification strategies to facilitate future investigation on structure-based methods for this task. Lastly, we showcase the powerful interpretability of structure-based methods through a case study on the Golgi apparatus, where we discover a decisive localization pattern $\alpha$-helix from attention mechanisms, demonstrating the potential for bridging the gap with intuitive biological interpretability and paving the way for data-driven discoveries in cell biology.
Tags
Links
- Source: https://arxiv.org/abs/2603.18571v1
- Canonical: https://arxiv.org/abs/2603.18571v1
Trouble viewing inline? Open PDF directly →
Full Text
116,463 characters extracted from source content.
Expand or collapse full text
Published as a conference paper at ICLR 2026 CAPSUL: A COMPREHENSIVE HUMAN PROTEIN BENCHMARK FOR SUBCELLULAR LOCALIZATION Yicheng Hu 1 , Xinyu Lin 2 , Shulin Li 3∗ , Wenjie Wang 1 , Fengbin Zhu 2† , Fuli Feng 1 1 University of Science and Technology of China, 2 National University of Singapore, 3 Tsinghua University ABSTRACT Subcellular localization is a crucial biological task for drug target identification and function annotation. Although it has been biologically realized that subcel- lular localization is closely associated with protein structure, no existing dataset offers comprehensive 3D structural information with detailed subcellular localiza- tion annotations, thus severely hindering the application of promising structure- based models on this task. To address this gap, we introduce a new benchmark called CAPSUL, a Comprehensive humAn Protein benchmark for SUbcellular Localization. It features a dataset that integrates diverse 3D structural represen- tations with fine-grained subcellular localization annotations carefully curated by domain experts. We evaluate this benchmark using a variety of state-of-the-art sequence-based and structure-based models, showcasing the importance of in- volving structural features in this task. Furthermore, we explore reweighting and single-label classification strategies to facilitate future investigation on structure- based methods for this task. Lastly, we showcase the powerful interpretability of structure-based methods through a case study on the Golgi apparatus, where we discover a decisive localization pattern α-helix from attention mechanisms, demonstrating the potential for bridging the gap with intuitive biological inter- pretability and paving the way for data-driven discoveries in cell biology. 1INTRODUCTION Understanding the subcellular localization of proteins is a fundamental question in cell biology, as a protein’s function is often tightly coupled to its spatial context within the cell (Scott et al., 2005). Localization information is essential for elucidating molecular mechanisms such as signal transduc- tion, metabolic regulation, and organelle-specific functions (Hung et al., 2017). It also provides a foundation for translational applications such as drug design (Hung et al., 2017; Rajendran et al., 2010). Recently, the data-driven AI approaches have emerged as a powerful paradigm for predicting whether or not a protein will be localized to a specific subcellular location. These methods substan- tially reduce the time and cost associated with traditional experimental techniques while holding promise for revealing novel biological patterns, thereby showcasing promising performance and at- tracting extensive research attention (Thumuluri et al., 2022; St ̈ ark et al., 2021; Almagro Armenteros et al., 2017; Kobayashi et al., 2022; Elnaggar et al., 2021). However, there remains a significant scarcity of high-quality datasets designed for this task. To the best of our knowledge, the only widely accepted dataset targeting this problem in the AI field is DeepLoc (Thumuluri et al., 2022; Almagro Armenteros et al., 2017), which contains the amino acid sequence information for each protein. DeepLoc has spurred the development of numerous sequence-based models for subcellular localization that infer localization solely from amino acid se- quences. Nevertheless, several studies have shown that spatial conformations play a critical role in determining subcellular localization patterns. For example, the nuclear localization signals of tran- scription factor NF-κB are conditionally exposed only under specific structural conformations (Lusk et al., 2007). This demonstrates that the 3D structures of proteins, as dynamic regulatory elements, are the key to governing their subcellular localization. ∗ Corresponding author. Email: lsl19@tsinghua.org.cn † Corresponding author. Email: zhfengbin@gmail.com 1 arXiv:2603.18571v1 [cs.AI] 19 Mar 2026 Published as a conference paper at ICLR 2026 To fully leverage protein structural data, recent research has developed structure-based protein rep- resentation models. Benefiting from the emergence of AlphaFold2 (Jumper et al., 2021), which offers reliable structural predictions for a vast number of proteins, the structure-based methods learn representations directly from the spatial geometry of proteins. Such approaches have demonstrated impressive performance across a range of tasks, including protein classification (Jing et al., 2020; Zhang et al., 2022; Fan et al., 2022) and protein generation (Dauparas et al., 2022; Watson et al., 2023), showcasing their ability to capture complex structural patterns beyond what sequence alone can provide. These successful implementations underscore the substantial potential of incorporating structural information into subcellular localization prediction frameworks. However, the existing subcellular localization datasets, such as DeepLoc, suffer from several lim- itations, which hinder the investigation of structure-based methods. Most notably, 1) they lack explicit protein 3D information, which is the key input to structure-based methods. Furthermore, 2) the current dataset typically uses coarse-grained compartment classifications, grouping subcellu- lar areas into broad categories (e.g., do not distinguish nuclear membrane and nucleoli in nucleus), which overlooks the unique localization characteristics and mechanisms associated with different organelles. Therefore, it leads to poor interpretability and great difficulty in discovering distinct patterns and underlying biological principles. To address these limitations, we aim to construct a human protein subcellular localization dataset that can facilitate research on structure-based methods for localization prediction and enable the discovery of more specific and biologically relevant localization patterns. Specifically, we have two considerations for the dataset: 1) Comprehensive 3D information, which seeks to enhance the comprehensiveness of the dataset by recording detailed localization data from different databases and integrating 3D structural information of proteins, thereby bringing convenience and providing a unified evaluation benchmark for structure-based prediction models within the community; 2) Fine-grained subcellular categorization, which aims to incorporate finer-grained localization la- bels with annotations based on biological empirical evidence. As such, researchers are allowed to investigate protein localization patterns at a more detailed and functionally meaningful level. To this end, we take the initiative of building a dataset called CAPSUL that simultaneously ful- fills the two considerations. Specifically, to obtain the 3D information, we leverage AlphaFold2 to extract the Cartesian coordinates of the Cα (alpha carbon) and utilize the FoldSeek to derive corresponding 3Di structural tokens for each protein, promoting structure understanding such as backbone conformation, folding patterns, and local structure. Moreover, to obtain comprehensive subcellular localization labels, we cross-reference each protein with annotation data from both the UniProt (Consortium, 2019) and Human Protein Atlas (HPA) (Thul et al., 2017) databases. Building upon the categories in the existing dataset DeepLoc, we further refine the subcellular area space by introducing 20 aggregated subcellular compartments, carefully curated and validated by domain ex- perts. We extend several state-of-the-art (SOTA) protein representation models to this downstream task and evaluate their performance on CAPSUL. To facilitate future research, we investigate several potential optimization strategies for structure-based model training and make innovative use of the attention mechanism to enhance the interpretability of protein subcellular localization patterns by integrating Transformer modules into existing models. Empirical results on CAPSUL validate the necessity of 3D information incorporation and the potential of leveraging structure-based methods for causal biology pattern discovery on the subcellular localization task. In summary, the contributions of this paper are threefold: • We represent the first systematic attempt to construct a human protein subcellular localization dataset with comprehensive 3D information, fine-grained categorization of cell compartments, and cross-referenced localization labels with experiment-level annotations. • We evaluate several SOTA baseline models on our proposed dataset CAPSUL, validating the positive influence of incorporating protein structural inputs. • We investigate various training strategies to facilitate future exploration and enhance the inter- pretability for subcellular localization tasks by introducing the attention mechanism. 2RELATED WORK Sequence-based protein representation learning. Due to the relative ease of modeling protein amino acid sequences, early protein representation learning efforts typically relied solely on one- 2 Published as a conference paper at ICLR 2026 dimensional sequence inputs. Examples include models based on CNN, LSTM, or ResNet archi- tectures (Shanehsazzadeh et al., 2020; Rao et al., 2019). Subsequently, Transformer-based models have demonstrated strong performance, especially after large-scale pretraining, achieving impres- sive results across a range of downstream tasks (Rives et al., 2019; Lin et al., 2022; Madani et al., 2023). In parallel, various self-supervised approaches have further enhanced the model’s ability to capture meaningful features from protein sequences without a vast number of annotations (Rives et al., 2019; Lin et al., 2023; Elnaggar et al., 2021; Lu et al., 2020; He et al., 2021). However, in the subcellular localization task, which is known to be closely linked to protein structure, sequence- only models fall short of capturing the full complexity of protein features. As a result, incorporating 3D structural information has become increasingly recognized as essential for achieving richer and more comprehensive protein representations. Structure-based protein representation learning. Efforts to model protein structures have been explored from multiple perspectives, including representations at the protein surface level, residue level, and atomic level. The protein language model also starts to consider structural information as input to enhance its understanding of proteins (Hayes et al., 2025). These approaches have achieved impressive results in tasks such as protein design, structure generation, and function pre- diction (Gligorijevi ́ c et al., 2021; Gainza et al., 2020; Hermosilla et al., 2020; Hsu et al., 2022). Among them, models based on Graph Convolutional Network (GCN) have demonstrated consis- tently strong performance across various downstream tasks, highlighting their ability to effectively capture and interpret structural information (Fan et al., 2022; Jing et al., 2020; Zhang et al., 2022). However, most of these models require atomic or residue-level coordinate inputs, which are often missing from current benchmark datasets. To address this gap, we aim to construct a dataset specif- ically for the task of subcellular localization that incorporates 3D structural information, facilitating both the application and evaluation of structure-based models. Subcellular localization dataset. Although many prestigious and task-specific protein benchmarks exist (Rao et al., 2019; Kryshtafovych et al., 2023), their lack of subcellular localization annotations makes them inapplicable on this downstream task. To the best of our knowledge, the only well- known dataset for subcellular localization originates from the training data used in DeepLoc (Thu- muluri et al., 2022). Building on this, the PEER framework established a benchmark to evaluate baseline models on that dataset (Xu et al., 2022). However, the absence of 3D structural information makes it impossible to assess the performance of structure-based models that have already shown significant promise. To address this gap, we aim to reorganize and enrich the existing dataset by incorporating high-quality 3D structural information alongside fine-grained subcellular localization annotations. We further evaluate a range of representative baseline models on this updated dataset, with the goal of establishing a leading benchmark for subcellular localization prediction. 3CAPSUL DATASET To construct the CAPSUL dataset that offers 1) diverse and accessible 3D structural information, and 2) both detailed and aggregated subcellular localization annotations, we follow a multi-step curation process, as illustrated in Figure 1. 3.1PROCESSING OF PROTEIN SEQUENCE AND STRUCTURE DATA Collection and filter of protein data. We first retrieve all predicted human protein structures from the AlphaFold2 database (Jumper et al., 2021; Varadi et al., 2024), totaling 20,504 unique proteins. To ensure data quality and relevance, we filter this set by retaining only proteins marked as active in the UniProt database (Consortium, 2019), one of the most comprehensive and authoritative protein databases with well-documented annotations, resulting in a refined set of 20,401 proteins. Removal of fragmented structure predictions. Among the refined set, AlphaFold2 typically adopts a sliding-window strategy to long protein sequences that segments the sequence with over- lapping fragments to predict protein structure. To avoid inconsistencies of predicted coordinates that may arise during the stitching of these fragmented protein structures, we exclude such proteins from the dataset. After this step, we obtain 20,181 proteins of high quality and good consistency. Extraction and preprocessing of protein features. We preserve the full PDB files for each pro- tein, the original files downloaded from AlphaFold, including the positions of backbone atoms, side 3 Published as a conference paper at ICLR 2026 Figure 1: Procedures of CAPSUL dataset construction, including 3 key steps: Step 1 extracts and filters the sequence and structure data for each high-quality protein from AlphaFold2; Step 2 collects the annotations from UniProt and HPA for the resulting proteins in Step 1; Step 3 merges the struc- ture data and the annotations for each protein, which consists of protein ID, localization annotations, amino acid sequence, sequence length, 3Di tokens, and Cα coordinates, etc. chains, and other relevant structural features essential for molecular modeling and analysis. The coordinates of Cα atoms are extracted, which are important components for protein structure un- derstanding. Furthermore, we employ the FoldSeek (Van Kempen et al., 2024) toolkit to tokenize the 3D structure of each amino acid. This provides a compact, informative structural representa- tion that supports rapid, accurate modeling while reducing computational overhead, which has been empirically justified as effective and widely adopted in recent studies (Su et al., 2023). Following the procedures above, we curate a dataset comprising 20,181 proteins, each labeled with amino acid sequence, Cα coordinates, and 3Di tokens sequence. For the next step, we append localization annotations to each protein. 3.2PROCESSING OF SUBCELLULAR LOCALIZATION ANNOTATIONS Acquisition of detailed subcellular localization annotations. Based on the obtained proteins above, we collect the corresponding detailed subcellular localization annotations for human proteins from both the UniProt and HPA databases. This detailed dataset provides high-resolution localiza- tion annotations on widely accepted subcellular compartments, which is vital to facilitate research into the specific localization patterns within distinct organelles. Fine-grained categorization. After that, we aggregate the dataset by adopting a refined catego- rization approach. Specifically, we consolidate the subcellular locations into 20 distinct categories inspired by DeepLoc’s and HPA’s subcellular localization classification scheme (Thumuluri et al., 2022; Thul et al., 2017), which is a fine-grained framework compared with DeepLoc’s ten-class categorization. Then, the sublocations of 20 categories are specified separately, so the various ter- minologies in different databases can align with 20 unified categorizations. The entire procedure was conducted in accordance with a well-established cell biology textbook (Alberts et al., 2022) and further verified by domain experts, with detailed categorization information available in Supp. A. Annotations of evidence level for localization data. To fulfill the various research demands for the reliability of localization labels, we further extract and consolidate annotations on the experimental evidence level. Specifically, for UniProt, each subcellular localization annotation is accompanied by an evidence code indicating the source of the localization label. Among them, the localization supported by experimental evidence (marked with the term ECO:0000269) is labeled as 1, indicating experimental validation. For the localization with other forms of evidence (e.g., non-traceable author statement evidence), the label 2 is assigned. The label 0 is assigned to the localizations without evidence annotations. Moreover, since HPA primarily relies on experimental data obtained through immunofluorescence and confocal microscopy (Thul et al., 2017), we assign label 1 to all annotated 4 Published as a conference paper at ICLR 2026 Table 1: Comparisons between existing datasets and CAPSUL. FearturesCategoization Experimental Annotation DatasetSequenceStructureAggregatedDetailed DeepLoc (Thumuluri et al., 2022)✓✗✓✗ setHARD (St ̈ ark et al., 2021)✓✗✓✗ CAPSUL✓ Table 2: Statistics of CAPSUL. Number of Proteins20,181Average Number of Annotations per Protein 2.51Max Number of Annotations for Protein 14Proportion of Experimental Annotations 0.857 Number of Annotations on: Nucleus7,590Cytosol5,386Golgi Apparatus1,881Peroxisome110 Nuclear Membrane452Cytoskeleton2,119Cell Membrane5,777Vesicle2,863 Nucleoli1,641Centrosome1,000Endosome687Primary Cilium983 Nucleoplasm6,786Mitochondria1,768Lipid Droplet94Secreted Proteins2,087 Cytoplasm6,613Endoplasmic Reticulum1,710Lysosome/Vacuole453Sperm652 localizations and label 0 to the localizations without evidence annotations. During the union of UniProt and HPA datasets, we prioritize annotations with experimental evidence when available. 3.3DATA MERGING After the separate processing of protein sequence and structure data, along with the subcellular localization annotations, we merge the data to include complete information. In Figure 1, we present a sample record in CAPSUL, which consists of protein ID, localization annotations, amino acid sequence, sequence length, 3Di tokens, and Cα coordinates, etc. 3.4DATASET ANALYSIS In summary, we construct a unified dataset comprising 20,181 proteins, each annotated with 20 sub- cellular localization labels. Our dataset CAPSUL provides a more comprehensive coverage com- pared to DeepLoc (Thumuluri et al., 2022) and setHARD (St ̈ ark et al., 2021) in terms of involved features, localization categorization, and experimental annotations, which is shown in Table 1. The dataset is randomly split into training, validation, and test sets in a 70%:15%:15% ratio for training and evaluation. We present a statistical analysis of numerical features of our dataset in Table 2. To ensure the high quality of our constructed CAPSUL dataset, we have incorporated three safe- guards 1 : 1) Reliable data sources: reliable protein structures predicted by AlphaFold2 were utilized in CAPSUL, with high accuracy, strong consistency, and incorporation of available experimen- tal data as templates in its prediction process (Jumper et al., 2021); the localization labels source UniProt, a world-leading database with the most comprehensive protein annotations from multiple resources, and HPA, a human-specific protein database offering high-resolution and experiment- validated data. 2) Strict validation and filtering: we perform a series of validation and filtering steps on human proteins to exclude fragmented AlphaFold structures, which could introduce in- consistent coordinate information, and to remove proteins annotated as inactive in UniProt, thereby ensuring the reliability of subcellular localization annotations; 3) Evidence-level support: we in- corporate annotations indicating whether experimental validation exists for the localization labels, thereby enhancing their credibility and catering to diverse research needs. 4EXPERIMENTS 4.1BASELINE MODELS To study how existing methods perform on our proposed dataset 2 , we evaluate 1) DeepLoc 2.1 (Ødum et al., 2024), one of the most well-known tools dedicated to subcellular localization. It leverages the pre-trained protein language model ESM-1b (Rives et al., 2021) and provides pre- dictions across ten subcellular compartments. Besides, we evaluate existing representative protein 1 For a detailed analysis of the data reliability in the dataset, please refer to Supp. B. 2 The detailed descriptions and hyperparameter settings of all baseline models are provided in Supp. C. 5 Published as a conference paper at ICLR 2026 representation methods for the subcellular localization task, including sequence-based and structure- based methods. Sequence-based models. Since existing sequence-based works are not specifically designed for subcellular tasks, we extend the widely adopted pre-trained protein language model 2) ESM-2 (650M parameters) (Lin et al., 2022) and its latest iteration, 3) ESM-C (600M parameters) (ESM Team, 2024). We adopt the sequence encoder module from existing methods to obtain protein rep- resentation, and extend it with a localization classifier, as detailed in the following. • Sequence Encoder. For each protein, we have its amino acid sequence represented as S = (s 1 , s 2 ,..., s n ) ∈R n×1 , where s i denotes the i-th residue and n is the length of the protein. We then apply the sequence encoder f seq (·) of existing work to obtain contextual embeddings, H = f seq (S), where H = (h 1 , h 2 ,..., h n ) ∈R n×d , and h is the per-residue embeddings of dimension d. To obtain a fixed-length representation for the entire protein, we apply mean pooling and generate a global representation ̄ h = 1 n P n i=1 h i , ̄ h∈R d . • Localization Classifier. To predict subcellular localization, we leverage an MLP classifier φ(·) on top of sequence encoder, i.e., ˆy = φ( ̄ h), where ˆy ∈R m is a multi-label prediction vector and m denotes the total number of predicted subcellular compartments. Structure-based models. We consider 4) CDConv (Fan et al., 2022) and 5) GearNet-Edge (Zhang et al., 2022), two representative GCN baselines in protein representation task. We adopt the GCN- based structure encoder and extend it with an additional Transformer encoder to enhance inter- pretability. We also evaluate 6) FoldSeek (Van Kempen et al., 2024), which leverages a pre-trained structure tokenizer to encode the 3D structural information of each residue into a sequence of struc- ture tokens. The outputs of the above models are then averaged and processed through a localization classifier for prediction. • Structure Encoder. We represent a protein’s 3D structure as a graph G = (V,E), where each node v i ∈ V corresponds to the i-th residue (typically using the Cα atom position), and edges (v i ,v j )∈ E are defined based on spatial or sequential adjacency. Each node v i is initialized with a feature vector x i ∈R d including its positional information. Then we employ different graph encoders to capture higher-order topological relationships and produce updated representations (h 1 ,..., h n ). The protein-level embedding is then obtained via global pooling ̄ h = 1 n P n i=1 h i . • Localization Classifier. We then obtain the final prediction ˆ y = φ ̄ h , as described above. Extension of structure-based models. We also extend three novel methods, 7) Graph Trans- former (Ramp ́ a ˇ sek et al., 2022), 8) Graph Mamba (Gu & Dao, 2023), and 9) Graph Diffu- sion (Yang et al., 2023) to this task. The Graph Transformer employs attention mechanisms over graph-structured data, enabling the model to effectively capture both local and global dependen- cies among residues. Graph Mamba, on the other hand, incorporates selective state space models into graph learning, which facilitates long-range information propagation with improved efficiency. For Graph Diffusion, it leverages diffusion processes over graph-structured data to propagate in- formation across nodes. To improve classification performance on minor categories, we further incorporate a contrastive loss mechanism into the CDConv model, i.e., 10) CDConv with Con- trastive Loss, aiming to enhance the similarity between representations of positive protein pairs and thereby encourage the learning of distinctive localization features. We also explore a fusion model that combines representative structure-based models with sequence-based pretrained protein language models, i.e., 11) ESM-C+CDConv Fusion Model, investigating both early and late fusion strategies to integrate structural information into large-scale sequence models. Optimization. To optimize the models, we adopt the Binary Cross Entropy (BCE) loss, defined asL BCE = − 1 m P m i=1 [y i log( ˆ y i ) + (1− y i ) log(1− ˆ y i )], where m is the number of classes, y i ∈ 0, 1 is the label for class i, and ˆ y i ∈ (0, 1) is the predicted probability. 4.2BENCHMARK OVERALL RESULTS AND DISCUSSION Given the class imbalance in each location (i.e., the proportion of proteins localized to each sub- cellular compartment is often small), we consider the widely used evaluation metrics in this task: Precision, Recall, and F1-score (Jiang et al., 2021; Thumuluri et al., 2022). In addition, we utilize micro-averaged and macro-averaged F1-score to evaluate the overall performance across different 6 Published as a conference paper at ICLR 2026 Table 3: Overall performance of sequence-based, structure-based methods on CAPSUL. Sequence-based MethodsStructure-based Methods Subcellular Locations DeepLoc 2.1 ESM-2 650M ESM-2 650M f ESM-C 600M ESM-C 600M f ESM-C 600M 0 FoldSeekCDConv t GearNet- Edge t F1-score Nucleus0.152-0.6090.6490.6480.5550.4840.6200.521 Nuclear Membrane/-------- Nucleoli/--0.0910.0390.024-0.1470.121 Nucleoplasm/-0.5620.6210.6230.5000.4330.5830.515 Cytoplasm0.154-0.2480.5360.5510.4380.1740.4830.495 Cytosol/--0.3920.3800.1690.0030.3530.385 Cytoskeleton/-0.0060.2510.2050.0480.0700.1350.228 Centrosome/--0.014----0.127 Mitochondria0.120-0.3170.5620.5440.099-0.4760.318 Endoplasmic Reticulum0.121--0.3510.3330.059-0.2920.279 Golgi Apparatus0.061--0.0990.027--0.0730.026 Cell Membrane0.142-0.5550.6310.6480.3720.3430.5620.556 Endosome/--0.018----0.067 Lipid Droplet/-------- Lysosome / Vacuole0.118-------0.073 Peroxisome0.131-------- Vesicle/--0.009-0.005-0.0270.068 Primary Cilium/--0.1640.112---0.147 Secreted Proteins0.191-0.7130.8260.7970.4330.3280.7670.687 Sperm/--0.0520.070---0.086 Micro Avg F1-score/-0.3750.4950.4920.3380.2480.4520.417 Macro Avg F1-score/-0.1500.2630.2490.1350.0920.2260.235 Micro Avg Precision/-0.6470.6900.6930.5980.6050.6320.546 Micro Avg Recall/-0.2640.3860.3820.2360.1560.3520.337 Extension of Structure-based Methods Subcellular Locations Graph Transformer Graph Mamba Graph Diffusion CDConv t with Contrastive Loss ESM-C+CDConv Early Fusion ESM-C+CDConv Late Fusion F1-score Nucleus0.5970.5590.6240.5920.6430.645 Nuclear Membrane-0.037---- Nucleoli0.2030.1680.0470.1400.1250.153 Nucleoplasm0.5520.5020.5780.5560.6430.617 Cytoplasm0.3930.4180.5030.4800.4550.515 Cytosol0.2480.4260.2880.2500.1570.370 Cytoskeleton0.0420.2700.0990.2430.1000.287 Centrosome-0.1810.014--0.037 Mitochondria0.4750.3410.3030.4680.5630.557 Endoplasmic Reticulum0.1840.0590.1500.3610.446- Golgi Apparatus0.0410.185-0.1560.249- Cell Membrane0.5470.5400.4960.5390.6290.673 Endosome-0.100-0.0340.054- Lipid Droplet------ Lysosome / Vacuole----0.026- Peroxisome------ Vesicle0.0440.1350.0180.0270.067- Primary Cilium0.0120.088-0.036-0.115 Secreted Proteins0.7050.5570.6230.7800.8190.725 Sperm0.0180.130-0.018-- Micro Avg F1-score0.4100.4110.4240.4350.4700.476 Macro Avg F1-score0.2030.2350.1870.2340.2490.235 Micro Avg Precision0.6370.4140.5960.6500.7100.634 Micro Avg Recall0.3020.4080.3290.3260.3510.381 f We finetune the pre-trained protein language model. t The original MLP is replaced by Transformer layers. 0 The parameters of ESM-C are initialized randomly. “/” indicates that DeepLoc 2.1 does not support prediction for that location, and therefore, average metrics are not considered in this case. “–” indicates that no prediction is made for that location. Bold value indicates the best results. Table 4: Ablation study of CDConv and GearNet-Edge to randomly sample Cα coordinates. CDConv t (random Cα coordinates) CDConv t GearNet-Edge t (random Cα coordinates) GearNet-Edge t Micro Avg F1-score0.3290.4520.3480.417 Micro Avg Precision0.5860.6320.4500.546 Micro Avg Recall0.2290.3520.2830.337 t The original MLP is replaced by Transformer layers. Bold value indicates the better result for each baseline. categories. The overall performance 3 of all baselines on our proposed dataset is presented in Table 3, from which we have the following observations: 3 The detailed results w.r.t. Precision and Recall, including other experimental results mentioned later in the main text, are provided in Supp. D. 7 Published as a conference paper at ICLR 2026 Large pre-training benefits sequence-based methods for subcellular location prediction. Among all sequence-based methods, ESM-C generally obtains higher F1-scores than ESM-2. We believe this is attributed to the extensive data and training compute used in the ESM-C pre-training, which facilitates a better representation of the protein’s sequence features. Similar observations are also seen in (Hayes et al., 2025). Besides, this hypothesis can be further confirmed by the signifi- cantly inferior performance of ESM-C 600M 0 , i.e., without pre-training, than the pre-trained ESM- C. On the other hand, it is expected that DeepLoc yields inferior performance due to its overlook of the fine-grained categorization during pre-training, which may result in its inability to sufficiently differentiate the representations of proteins in multi-label classification tasks (Hong et al., 2023). This further validates the necessity of detailed categorizations of subcellular locations in CAPSUL. The 3D structure is essential for subcellular localization task. Despite that structure-based meth- ods slightly fall behind the pre-trained ESM-C, both CDConv and GearNet-Edge outperform the ESM-C 600M 0 in most cases. Also, a group of ablation studies is conducted on CDConv and GearNet-Edge, with coordinates randomly sampled from each protein’s spatial range. As shown in Table 4, randomly sampling the input of protein 3D structural data leads to a significant drop in model performance. These two results validate that structural information plays a decisive role in determining subcellular localization. Besides, CDConv demonstrates the strongest overall perfor- mance among the structure-based models, justifying the effectiveness of relative distance and the dynamic radius for convolution. Nevertheless, the inferior performance of FoldSeek may be due to the lack of sequence information and its coarse tokenization of structural information. The models generally demonstrate better performance on subcellular locations with larger lo- calization sample sizes. For classes with a large number of localization samples (e.g., nucleus), most models tend to demonstrate relatively strong predictive performance. In contrast, for under- represented classes (e.g., lipid droplet), the prediction performance is generally poor, with some classes even failing to produce any correctly identified proteins. This is a common outcome in im- balanced multi-label classification tasks, as the standard BCE loss tends to neglect fewer-number labels. Additionally, potential conflicts among multiple optimization targets may further exacerbate this issue. To address these challenges, we conduct in-depth analysis in Section 4.3.1 and 4.3.2, ex- ploring strategies such as reweighting and single-label classification to mitigate the effects of class imbalance and task conflict. Structure-based models showcase their potential to capture non-trivial patterns for subcel- lular locations with few samples. Graph Mamba and GearNet-Edge tend to perform better on certain classes with smaller localization sample sizes compared with sequence-based models. We believe that this is because of the relational message passing layer adopted in them, which uniquely models different spatial interactions among residues. This demonstrates that structure-based mod- els showcase potential to identify specific structural features that are indicative of localization to a particular organelle, thus achieving a notably good performance. Further investigation on the pat- terns with intuitive biological interpretability captured by the structure-based model can be found in Section 4.3.3. Contrastive learning and fusion strategies demonstrate strong potential on the baseline mod- els. We observe that the introduction of contrastive loss to CDConv improves performance on several minority classes (i.e., the notable F1-score improvements on macro-average level and certain cate- gories such as Endosome, Primary Cilium, and Golgi Apparatus compared to the original CDConv). We attribute this to the contrastive learning paradigm, which encourages the model to maximize embedding similarity for positive pairs and to capture shared characteristics within minority-class positive samples through the contrastive objective. Also, although the fusion model slightly under- performs ESM-C on average metrics, it achieves the best performance across multiple subcellular compartments among all baselines, highlighting the considerable potential of integrating protein structural information into sequence-based protein language models. 4.3IN-DEPTH ANALYSIS 4.3.1PROTEIN IMBALANCE MITIGATION VIA REWEIGHTING Reweighting Schemes. In this task, for each subcellular location, the number of positive samples (i.e., proteins localized to that compartment) is substantially smaller than the number of negative samples (i.e., proteins not localized to that compartment). Reweighting is a widely used strategy to 8 Published as a conference paper at ICLR 2026 Table 5: Performance of ESM-C 600M, CDConv, and GearNet-Edge with reweighting scheme. Subcellular Locations ESM-C 600M CDConv t GearNet- Edge t Subcellular Locations ESM-C 600M CDConv t GearNet- Edge t F1-scoreF1-score Nucleus0.6300.6250.618Endosome-0.1140.150 Nuclear Membrane-0.0620.058Lipid Droplet0.2350.0230.111 Nucleoli-0.1880.224Lysosome/Vacuole-0.1750.111 Nucleoplasm0.5760.6070.574Peroxisome0.1900.0720.108 Cytoplasm0.5000.5820.544Vesicle-0.2880.281 Cytosol0.1330.4950.484Primary Cilium0.0240.1670.176 Cytoskeleton0.0830.2920.294Secreted Proteins0.7780.5640.614 Centrosome-0.1600.175Sperm-0.1200.125 Mitochondria0.4810.2970.313 Endoplasmic Reticulum-0.3080.345 Golgi Apparatus-0.2460.238Micro Avg F1-score0.4290.3810.453 Cell Membrane0.5660.5600.536Macro Avg F1-score0.2100.2970.304 t The original MLP is replaced by Transformer layers. “–” indicates that no prediction is made for that location. Bold value indicates that it improves compared with the result without reweighting. Table 6: Performance of ESM-C 600M, CDConv, and GearNet-Edge with single-label classification. Subcellular Locations ESM-C 600M CDConv t GearNet- Edge t Subcellular Locations ESM-C 600M CDConv t GearNet- Edge t F1-scoreF1-score Nuclear Membrane-0.0520.042Lysosome/Vacuole0.115-0.162 Nucleoli0.2670.1510.228Peroxisome0.054-0.023 Centrosome0.1840.0890.167Vesicle0.0680.2300.268 Golgi Apparatus0.2800.1140.210Primary Cilium0.2530.0970.171 Endosome0.1670.0490.126Sperm0.1590.0680.117 Lipid Droplet0.021-0.051 t The original MLP is replaced by Transformer layers. “–” indicates that no prediction is made for that location. Bold value indicates that it improves compared with the result from multi-label classification. address class imbalance by reducing the bias toward majority classes. Inspired by previous work on class-level reweighting, we evaluate three reweighting schemes. 1) Inverse frequency reweight- ing (Cao et al., 2019), i.e., w c = 1 f c . 2) Log-inverse frequency reweighting (Cui et al., 2019), i.e., w c = 1 log(1+f c ) . 3) Focal loss (Lin et al., 2017), which is defined as L c =−w c · X i [y ic · (1− ˆy ic ) γ log(ˆy ic ) + (1− y ic )· ˆy γ ic log(1− ˆy ic )], where f c is the frequency of positive samples in class c, w c is the computed class-specific weight, y ic ∈ 0, 1 denotes the ground truth label for sample i and class c, ˆy ic ∈ (0, 1) is the predicted probability, and γ is the focusing parameter. It deserves attention that the w c in the focal loss strategy is chosen from either inverse or log-inverse frequency weight. Results. We apply the three reweighting schemes on three competitive models (ESM-C, CDConv, and GearNet-Edge) and report the best results for each model in Table 5. From the results, we ob- serve that the two structure-based baseline models exhibit substantial improvements under reweight- ing strategies, especially for the higher Precision across underrepresented categories. In particular, CDConv and GearNet-Edge successfully identify positive instances for every class. These find- ings highlight that reweighting can significantly enhance model performance on minority classes, especially for structure-based models. 4.3.2SINGLE-LABEL CLASSIFICATION To explore how different methods perform on each subcellular location respectively, we adopt the single-label setting, aiming to mitigate the potential conflict between optimization across different classes. In this setting, we train separate binary classifiers for each subcellular localization cate- gory with ESM-C, CDConv, and GearNet-Edge. We apply this single-label prediction framework specifically to those subcellular localization classes where the F1-score of at least one of the models (ESM-C, CDConv, or GearNet-Edge) is lower than 0.1. Our goal is to shift the model’s attention toward underrepresented classes and improve the predictive influence of positive samples. From the results in Table 6, we observe 1) notable improvements in the prediction performance of previously underperforming classes, particularly for GearNet-Edge. However, 2) ESM-C and CDConv still fail to generate any predictions for a few categories, primarily due to the extremely low proportion of positive samples (ranging from 0.5% to 3%). Given that such severe class imbalance is a common challenge in subcellular localization tasks, we consider the single-label prediction 9 Published as a conference paper at ICLR 2026 Figure 2: Visualization of the top 20 attention-scored residues of the three representative proteins. strategy a promising and practical solution. Moreover, this approach lays the groundwork for future research focused on identifying localization patterns specific to individual subcellular compartments. 4.3.3BIOLOGICAL INTERPRETABILITY We analyze a CDConv model on Golgi apparatus prediction with an exceptional precision of 100%. Specifically, with our novel attempt of the Transformer module extended to the GCN-based models, we identify and visualize the tokens (i.e., residues) that receive the 20 highest attention weights in Figure 2, offering insights into which structure the model considers most decisive for subcellular lo- calization. We find that the model consistently highlights similar α-helix spatial conformation, such as residues 8-27 of MFNG, residues 24-45 of B3GALT2, and residues 273-292 at the C-terminus of GIMAP1. Remarkably, these findings show strong concordance with prior experimental evi- dence (Paulson & Colley, 1989; Linstedt et al., 1995). It is highlighted that despite significant se- quence divergence, the model specifically focuses on α-helix transmembrane domains (20-30 amino acids in length) that maintain consistent topological orientations across all targets. Recent studies have demonstrated that the topological conformation of transmembrane domains can influence Golgi localization by regulating electrostatic potential gradients in transmembrane regions and lipid mem- brane anchoring efficiency (Cosson et al., 2013; Hanulova & Weiss, 2012; Bian et al., 2024). This evidence not only confirms the model’s capability for structural pattern recognition beyond sequence similarity but also provides theoretical support for its structural identification mechanisms. 5CONCLUSION AND FUTURE WORK We pointed out the crucial importance of constructing a subcellular localization benchmark with protein 3D information to facilitate the investigation of structure-based models for the subcellu- lar localization task. To achieve this, we constructed a benchmark called CAPSUL that contains comprehensive structural information and fine-grained annotations of 20 categories of subcellular compartments with biological experiment evidence labels. Based on CAPSUL, we evaluated SOTA sequence-based and structure-based models as well as their feasible optimization strategies, demon- strating the effectiveness of incorporating protein structural information. Moreover, a case study on Golgi apparatus validates the biology-aligned interpretability of structure-based models trained on a specific fine-grained subcellular location, supported by CAPSUL. This work proposes a compre- hensive human protein benchmark with 3D information and fine-grained annotations for subcellular localization. Based on CAPSUL, we highlight several research directions that are worth future ex- ploration: 1) To fully leverage structural information, aligning or disentangling the understanding across different dimensions (i.e., amino acid sequence, Cα, and 3Di) specifically for subcellular localization is a promising direction. 2) Causal discovery on the relationship between 3D structure and subcellular localization is worthwhile to be explored on CAPSUL, with the goal of establishing direct links to underlying biological principles. ACKNOWLEDGMENT Shulin Li was supported by China Postdoctoral Science Foundation (grant number BX20240186 and 2024M761616) and the Shuimu Tsinghua Scholar Program. 10 Published as a conference paper at ICLR 2026 ETHICS STATEMENT This research presents a dataset and benchmark for protein subcellular localization prediction using AI methods. We confirm that our work raises no ethical concerns as it involves only the analysis of publicly available protein data, with no human subjects, animal experiments, or biological in- terventions. We have fully considered the potential societal impacts and do not foresee any direct, immediate, or negative consequences. We are committed to the ethical dissemination of our findings and encourage their responsible use. REPRODUCIBILITY STATEMENT All the results in this work are reproducible. The access to the necessary code and complete dataset can be found in Supp. L. We discuss the experimental details in Supp. C, including implementa- tion details such as the hyperparameters chosen for each experiment, to help reproduce our results. Additionally, further experimental results, detailed dataset interpretations, and usage guidelines are provided in Supp. E, F, and J to facilitate better understanding and utilization of our dataset. REFERENCES Bruce Alberts, Rebecca Heald, Alexander Johnson, David Morgan, Martin Raff, Keith Roberts, and Peter Walter. Molecular biology of the cell: Seventh edition. Norton and Company, 2022. Jos ́ e Juan Almagro Armenteros, Casper Kaae Sønderby, Søren Kaae Sønderby, Henrik Nielsen, and Ole Winther. Deeploc: prediction of protein subcellular localization using deep learning. Bioinformatics, 33(21):3387–3395, 2017. T. G. Ashlin, N. J. Blunsom, and S. Cockcroft. Courier service for phosphatidylinositol: Pitps deliver on demand. Biochimica et Biophysica Acta (BBA) - Molecular and Cell Biology of Lipids, 1866(9):158985, 2021. Claudie Bian, Anna Marchetti, Marco Dias, Jackie Perrin, and Pierre Cosson. Short transmembrane domains target type i proteins to the golgi apparatus and type i proteins to the endoplasmic reticulum. Journal of Cell Science, 137(15), 2024. Kaidi Cao, Colin Wei, Adrien Gaidon, Nikos Arechiga, and Tengyu Ma. Learning imbalanced datasets with label-distribution-aware margin loss. Advances in neural information processing systems, 32, 2019. UniProt Consortium. Uniprot: a worldwide hub of protein knowledge. Nucleic acids research, 47 (D1):D506–D515, 2019. Pierre Cosson, Jackie Perrin, and Juan S Bonifacino. Anchors aweigh: protein localization and transport mediated by transmembrane domains. Trends in cell biology, 23(10):511–517, 2013. Yin Cui, Menglin Jia, Tsung-Yi Lin, Yang Song, and Serge Belongie. Class-balanced loss based on effective number of samples. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 9268–9277, 2019. Justas Dauparas, Ivan Anishchenko, Nathaniel Bennett, Hua Bai, Robert J Ragotte, Lukas F Milles, Basile IM Wicky, Alexis Courbet, Rob J de Haas, Neville Bethel, et al. Robust deep learning– based protein sequence design using proteinmpnn. Science, 378(6615):49–56, 2022. Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, et al. Prottrans: towards crack- ing the language of life’s code through self-supervised learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44:7112–7127, 2021. ESM Team. Esm cambrian: Revealing the mysteries of proteins with unsupervised learning, 2024. URL https://evolutionaryscale.ai/blog/esm-cambrian. 11 Published as a conference paper at ICLR 2026 Hehe Fan, Zhangyang Wang, Yi Yang, and Mohan Kankanhalli. Continuous-discrete convolution for geometry-sequence modeling in proteins. In The Eleventh International Conference on Learning Representations, 2022. Pablo Gainza, Freyr Sverrisson, Frederico Monti, Emanuele Rodola, Davide Boscaini, Michael M Bronstein, and Bruno E Correia. Deciphering interaction fingerprints from protein molecular surfaces using geometric deep learning. Nature Methods, 17(2):184–192, 2020. Vladimir Gligorijevi ́ c, P Douglas Renfrew, Tomasz Kosciolek, Julia Koehler Leman, Daniel Beren- berg, Tommi Vatanen, Chris Chandler, Bryn C Taylor, Ian M Fisk, Hera Vlamakis, et al. Structure- based protein function prediction using graph convolutional networks. Nature communications, 12(1):3168, 2021. Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023. Maria Hanulova and Matthias Weiss. Membrane-mediated interactions–a physico-chemical basis for protein sorting. Molecular Membrane Biology, 29(5):177–185, 2012. Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q Tran, Jonathan Deaton, Marius Wiggert, et al. Simulating 500 million years of evolution with a language model. Science, p. eads0018, 2025. Liang He, Shizhuo Zhang, Lijun Wu, Huanhuan Xia, Fusong Ju, He Zhang, Siyuan Liu, Yingce Xia, Jianwei Zhu, Pan Deng, et al. Pre-training co-evolutionary protein representation via a pairwise masked language model. arXiv preprint arXiv:2110.15527, 2021. Pedro Hermosilla, Marco Sch ̈ afer, Mat ˇ ej Lang, Gloria Fackelmann, Pere Pau V ́ azquez, Barbora Kozl ́ ıkov ́ a, Michael Krone, Tobias Ritschel, and Timo Ropinski. Intrinsic-extrinsic convolution and pooling for learning on 3d protein structures. arXiv preprint arXiv:2007.06252, 2020. Guan Zhe Hong, Yin Cui, Ariel Fuxman, Stanley H Chan, and Enming Luo. Towards understanding the effect of pretraining label granularity. arXiv preprint arXiv:2303.16887, 2023. Chloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin, Brian Hie, Tom Sercu, Adam Lerer, and Alexan- der Rives. Learning inverse folding from millions of predicted structures. In International con- ference on machine learning, p. 8946–8970. PMLR, 2022. Victoria Hung, Stephanie S Lam, Namrata D Udeshi, Tanya Svinkina, Gaelen Guzman, Vamsi K Mootha, Steven A Carr, and Alice Y Ting. Proteomic mapping of cytosol-facing outer mito- chondrial and er membranes in living human cells by proximity biotinylation. elife, 6:e24463, 2017. Alexander M Ille, Christopher Markosian, Stephen K Burley, Renata Pasqualini, and Wadih Arap. Human protein interactome structure prediction at scale with boltz-2. bioRxiv, p. 2025–07, 2025. Yuexu Jiang, Duolin Wang, Yifu Yao, Holger Eubel, Patrick K ̈ unzler, Ian Max Møller, and Dong Xu. Mulocdeep: a deep-learning framework for protein subcellular and suborganellar localization prediction with residue-level interpretation. Computational and structural biotechnology journal, 19:4825–4839, 2021. Bowen Jing, Stephan Eismann, Patricia Suriana, Raphael JL Townshend, and Ron Dror. Learning from protein structure with geometric vector perceptrons. arXiv preprint arXiv:2009.01411, 2020. John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin ˇ Z ́ ıdek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. nature, 596(7873):583–589, 2021. Hirofumi Kobayashi, Keith C Cheveralls, Manuel D Leonetti, and Loic A Royer. Self-supervised deep learning encodes high-resolution features of protein subcellular localization. Nature meth- ods, 19(8):995–1003, 2022. 12 Published as a conference paper at ICLR 2026 Andriy Kryshtafovych, Torsten Schwede, Maya Topf, Krzysztof Fidelis, and John Moult. Critical assessment of methods of protein structure prediction (casp)—round xv. Proteins: Structure, Function, and Bioinformatics, 91(12):1539–1549, 2023. Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Doll ́ ar. Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision, p. 2980–2988, 2017. Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Sal Candido, et al. Language models of protein sequences at the scale of evolution enable accurate structure prediction. bioRxiv, 2022. Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637):1123–1130, 2023. AD Linstedt, M Foguet, M Renz, HP Seelig, BS Glick, and HP Hauri. A c-terminally-anchored golgi protein is inserted into the endoplasmic reticulum and then transported to the golgi apparatus. Proceedings of the National Academy of Sciences, 92(11):5102–5105, 1995. Amy X Lu, Haoran Zhang, Marzyeh Ghassemi, and Alan Moses. Self-supervised contrastive learn- ing of protein representations by mutual information maximization. BioRxiv, p. 2020–09, 2020. C Patrick Lusk, G ̈ unter Blobel, and Megan C King. Highway to the inner nuclear membrane: rules for the road. Nature Reviews Molecular Cell Biology, 8(5):414–420, 2007. Ali Madani, Ben Krause, Eric R Greene, Subu Subramanian, Benjamin P Mohr, James M Holton, Jose Luis Olmos Jr, Caiming Xiong, Zachary Z Sun, Richard Socher, et al. Large language models generate functional protein sequences across diverse families. Nature biotechnology, 41 (8):1099–1106, 2023. Marius Thrane Ødum, Felix Teufel, Vineet Thumuluri, Jos ́ e Juan Almagro Armenteros, Alexan- der Rosenberg Johansen, Ole Winther, and Henrik Nielsen. Deeploc 2.1: multi-label membrane protein type prediction using protein language models. Nucleic Acids Research, 52(W1):W215– W220, 2024. Saro Passaro, Gabriele Corso, Jeremy Wohlwend, Mateo Reveiz, Stephan Thaler, Vignesh Ram Somnath, Noah Getz, Tally Portnoi, Julien Roy, Hannes Stark, et al. Boltz-2: Towards accurate and efficient binding affinity prediction. BioRxiv, 2025. James C Paulson and Karen J Colley. Glycosyltransferases: structure, localization, and control of cell type-specific glycosylation. Journal of Biological Chemistry, 264(30):17615–17618, 1989. Lawrence Rajendran, Hans-Joachim Kn ̈ olker, and Kai Simons. Subcellular targeting strategies for drug design and delivery. Nature reviews Drug discovery, 9(1):29–42, 2010. Ladislav Ramp ́ a ˇ sek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu, Guy Wolf, and Do- minique Beaini. Recipe for a general, powerful, scalable graph transformer. Advances in Neural Information Processing Systems, 35:14501–14515, 2022. Roshan Rao, Nicholas Bhattacharya, Neil Thomas, Yan Duan, Peter Chen, John Canny, Pieter Abbeel, and Yun Song. Evaluating protein transfer learning with tape. Advances in neural infor- mation processing systems, 32, 2019. Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C. Lawrence Zitnick, Jerry Ma, and Rob Fergus. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. PNAS, 2019. doi: 10.1101/622803. URL https://w.biorxiv.org/content/10.1101/622803v4. Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences, 118(15):e2016239118, 2021. 13 Published as a conference paper at ICLR 2026 Michelle S Scott, Sara J Calafell, David Y Thomas, and Michael T Hallett. Refining protein subcel- lular localization. PLoS computational biology, 1(6):e66, 2005. Amir Shanehsazzadeh, David Belanger, and David Dohan. Is transfer learning necessary for protein landscape prediction? arXiv preprint arXiv:2011.03443, 2020. S. Sohn, M. K. Joe, T. E. Kim, et al. Dual localization of wild-type myocilin in the endoplasmic reticulum and extracellular compartment likely occurs due to its incomplete secretion. Molecular Vision, 15:545, 2009. Hannes St ̈ ark, Christian Dallago, Michael Heinzinger, and Burkhard Rost. Light attention predicts protein location from the language of life. Bioinformatics Advances, 1(1):vbab035, 2021. Jin Su, Chenchen Han, Yuyang Zhou, Junjie Shan, Xibin Zhou, and Fajie Yuan. Saprot: Protein language modeling with structure-aware vocabulary. bioRxiv, p. 2023–10, 2023. Peter J Thul, Lovisa ̊ Akesson, Mikaela Wiking, Diana Mahdessian, Aikaterini Geladaki, Hammou Ait Blal, Tove Alm, Anna Asplund, Lars Bj ̈ ork, Lisa M Breckels, et al. A subcellular map of the human proteome. Science, 356(6340):eaal3321, 2017. Vineet Thumuluri, Jos ́ e Juan Almagro Armenteros, Alexander Rosenberg Johansen, Henrik Nielsen, and Ole Winther. Deeploc 2.0: multi-label subcellular localization prediction using protein lan- guage models. Nucleic acids research, 50(W1):W228–W234, 2022. Michel Van Kempen, Stephanie S Kim, Charlotte Tumescheit, Milot Mirdita, Jeongjae Lee, Cameron LM Gilchrist, Johannes S ̈ oding, and Martin Steinegger. Fast and accurate protein struc- ture search with foldseek. Nature biotechnology, 42(2):243–246, 2024. Mihaly Varadi, Damian Bertoni, Paulyna Magana, Urmila Paramval, Ivanna Pidruchna, Malarvizhi Radhakrishnan, Maxim Tsenkov, Sreenath Nair, Milot Mirdita, Jingi Yeo, et al. Alphafold protein structure database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic acids research, 52(D1):D368–D375, 2024. Joseph L Watson, David Juergens, Nathaniel R Bennett, Brian L Trippe, Jason Yim, Helen E Eise- nach, Woody Ahern, Andrew J Borst, Robert J Ragotte, Lukas F Milles, et al. De novo design of protein structure and function with rfdiffusion. Nature, 620(7976):1089–1100, 2023. Jeremy Wohlwend, Gabriele Corso, Saro Passaro, Noah Getz, Mateo Reveiz, Ken Leidal, Wojtek Swiderski, Liam Atkinson, Tally Portnoi, Itamar Chinn, et al. Boltz-1 democratizing biomolecular interaction modeling. BioRxiv, p. 2024–11, 2025. Minghao Xu, Zuobai Zhang, Jiarui Lu, Zhaocheng Zhu, Yangtian Zhang, Ma Chang, Runcheng Liu, and Jian Tang. Peer: a comprehensive and multi-task benchmark for protein sequence understand- ing. Advances in Neural Information Processing Systems, 35:35156–35173, 2022. R Yang, Y Yang, F Zhou, et al. Directional diffusion models for graph representation learning. Advances in Neural Information Processing Systems, 36:32720–32731, 2023. Zuobai Zhang, Minghao Xu, Arian Jamasb, Vijil Chenthamarakshan, Aurelie Lozano, Payel Das, and Jian Tang. Protein representation learning by geometric structure pretraining. arXiv preprint arXiv:2203.06125, 2022. 14 Published as a conference paper at ICLR 2026 Supplementary Material for CAPSUL: A Comprehensive Human Protein Benchmark for Subcellular Localization ADATASET CONSTRUCTION A.1SUBCELLULAR LOCATION CATEGORIZATION AND TERMINOLOGY MAPPING To facilitate model classification, we first categorize the detailed subcellular localizations of pro- teins. Existing datasets often use coarse-grained classifications (e.g., DeepLoc categorizes sub- cellular locations into 10 broad classes). However, since each subcellular compartment typically follows distinct localization patterns, such coarse categorizations can hinder the model’s ability to capture consistent intra-class features, ultimately leading to reduced prediction accuracy. Moreover, coarse-grained classification also hinders researchers from exploring localization mechanisms spe- cific to finer subcellular compartments. Inspired by the subcellular location categories in HPA and DeepLoc, we propose a finer-grained classification scheme consisting of 20 subcellular categories. Notably, ”Nucleus” and ”Cytoplasm” categories serve as umbrella terms for several finer locations to ensure compatibility with DeepLoc during evaluation. When aligning protein localization annotations from the UniProt and HPA databases to our refined categorization, we observe inconsistencies in terminology (e.g., ”Cell Membrane” in UniProt versus ”Plasma Membrane” in HPA). To resolve such discrepancies, we refer to the prestigious textbook Molecular Biology of the Cell (7th Edition) (Alberts et al., 2022) and create a unified mapping, as shown in Table 7, which allows for consistent categorization across the two databases. Domain experts were extensively engaged to ensure and validate the accuracy of the classification standards and data alignment procedures. We invited cell biologists from several prestigious uni- versities and research institutes to review and revise the dataset, which ensures that CAPSUL is firmly grounded in cell biology. All of them have over eight years of research experience in their field. They are rigorously involved throughout the entire process, including 1) curating authoritative datasets, 2) determining primary subcellular localizations, and 3) validating the biological plausibil- ity of localization assignments. Through the above processes, we have established a fine-grained subcellular localization classifica- tion standard and successfully unified annotations from multiple databases under a unified labeling framework. A.2DATASET SPLITS To construct separate datasets for training, validating, and testing, we randomly split the original dataset into three subsets in a 70%: 15%: 15% ratio, each containing 14,126, 3,027, 3,028 pro- teins. The partitioning of different protein data used in our experiments is also available in the CAPSUL dataset. The number of labels for each subcellular location in three subsets is shown in Table 8. Although the data is randomly assigned to different subsets, we have verified the distribu- tion characteristics among classes to maintain a similar proportional relationship, ensuring balance and representativeness across the subsets. BDATASET RELIABILITY B.1OVERVIEW OF DATA SOURCES In Section 3, we provide a detailed description of the data preprocessing procedures implemented to ensure the high quality of CAPSUL. Here, we would like to emphasize that the data sources themselves are highly reliable. Specifically, the protein-related data used in this study were primarily obtained from the following databases: AlphaFold. AlphaFold provides protein structural data in CAPSUL. 1) AlphaFold has already in- corporated experimentally resolved structures of proteins as templates during its prediction pro- cess (Jumper et al., 2021). AlphaFold explicitly describes how its pipeline automatically searches 15 Published as a conference paper at ICLR 2026 Table 7: Categorization of CAPSUL and terminology mapping between HPA and Uniprot. 20 fine-grained categoriesHPAUniProt Nucleus Nuclear MembraneNuclear membraneNucleus membrane, Nucleus envelope, Nucleus inner membrane, Nucleus outer membrane NucleoliNucleoli, Nucleoli fibrillar center, Nucleoli rim Nucleolus NucleoplasmKinetochore, Mitotic chromosome, Nuclear bodies, Nuclear speckles, Nucleoplasm Nucleus matrix, Nucleus lamina, Chromosome, Nucleus speckle Cytoplasm CytosolAggresome, Cytoplasmic bodies, Cytosol, Rods Rings Cytosol CytoskeletonActin filaments, Cleavage furrow, Focal adhesion sites, Cytokinetic bridge, Microtubule ends, Microtubules, Midbody, Midbody ring, Mitotic spindle, Intermediate filaments Cytoskeleton CentrosomeCentriolar satellite, CentrosomeCentrosome MitochondriaMitochondriaMitochondrion, Mitochondrion envelop, Mitochondrion inner membrane, Mitochondrion outer membrane, Mitochondrion membrane, Mitochondrion matrix, Mitochondrion intermembrane space Endoplasmic ReticulumEndoplasmic reticulumEndoplasmic reticulum, Endoplasmic reticulum membrane, Endoplasmic reticulum lumen, Microsome, Rough endoplasmic reticulum, Smooth endoplasmic reticulum, Sarcoplasmic reticulum Golgi ApparatusGolgi apparatusGolgi apparatus, Golgi apparatus membrane, Golgi apparatus lumen Cell MembraneCell Junctions, Plasma membraneCell membrane, Apical cell membrane, Apicolateral cell membrane, Basal cell membrane, Basolateral cell membrane, Lateral cell membrane, Cell projection EndosomeEndosomesEndosome Lipid DropletLipid dropletsLipid droplet Lysosome/VacuoleLysosomesLysosome, Vacuole, Vacuole lumen, Vacuole membrane, Lysosome lumen, Lysosome membrane PeroxisomePeroxisomesPeroxisome, Peroxisome matrix, Peroxisome membrane VesicleVesiclesVesicle Primary CiliumBasal body, Primary cilium, Primary cilium tip, Primary cilium transition zone Cilium Secreted ProteinsSecreted ProteinsSecreted SpermAcrosome, Annulus, Calyx, Connecting piece, End piece, Equatorial segment, Flagellar centriole, Mid piece, Perinuclear theca, Principal piece Acrosome, Calyx, Perinuclear theca 16 Published as a conference paper at ICLR 2026 Table 8: Label counts for training, validation, and test set of CAPSUL. Subcellular Locations Counts Training SetValidation SetTest SetSum Nucleus5,3121,1281,1507,590 Nuclear Membrane3136376452 Nucleoli1,1432492491,641 Nucleoplasm4,7511,0071,0286,786 Cytoplasm4,6529849776,613 Cytosol3,7878117885,386 Cytoskeleton1,4993023182,119 Centrosome7131401471,000 Mitochondria1,2472592621,768 Endoplasmic Reticulum1,1462752891,710 Golgi Apparatus1,3232712871,811 Cell Membrane4,0228638925,777 Endosome466113108687 Lipid Droplet63161594 Lysosome/Vacuole3136575453 Peroxisome712019110 Vesicle2,0194044402,863 Primary Cilium699123161983 Secreted Proteins1,4773172932,087 Sperm44499109652 the PDB for experimentally resolved structures, selecting up to four structural templates, and maps atom coordinates from those templates to the target sequence during inference. These coordinates are used as template inputs alongside MSA-based evolutionary information, enabling AlphaFold to leverage high-quality experimental structural data in its predictions. 2) AlphaFold-predicted struc- tures have been demonstrated to achieve exceptionally high accuracy, competitive with experimen- tal data. AlphaFold was entered for CASP14, and shows that it achieves accuracy competitive with experiment in a majority of cases. Specifically, the median backbone accuracy of its predictions is 0.96 ̊ A r.m.s.d. 95 (Cα root-mean-square deviation at 95% residue coverage), which is often within the margin of error of experimental structures (Jumper et al., 2021). 3) AlphaFold provides full- length protein structures containing complete structural information, which minimizes the potential negative influence of structural variability caused by different versions of experimental protein data. This choice allows us to maintain a high level of consistency across the CAPSUL dataset. UniProt. UniProt provides protein localization annotation and evidence-level annotations in CAP- SUL. UniProt serves as one of the most authoritative and widely used protein knowledge bases, integrating sequence, functional, and localization information across a broad spectrum of species. In particular, the manually curated Swiss-Prot section is recognized for its rigorous curation stan- dards, where annotations (including subcellular localization annotations) are derived from authori- tative experimental studies and peer-reviewed literature, complemented by computational analyses and homology-based inferences. Each localization entry is systematically annotated with evidence codes that explicitly denote whether the information originates from direct experimental validation, literature reports, or computational prediction, thereby providing transparency and traceability of the data source. This evidence-based framework ensures that localization annotations are not only comprehensive but also of consistently high quality. Human Protein Atlas (HPA). HPA provides protein localization annotation and subcellular cat- egories reference in CAPSUL. HPA provides a unique and experimentally grounded resource for human protein subcellular localization. Its Subcellular Atlas is built upon systematic immunoflu- orescence imaging combined with antibody-based profiling in multiple well-characterized human cell lines. This approach allows direct visualization of protein distribution within distinct subcel- lular compartments, thereby offering cell-type-specific and high-resolution localization evidence. These measures substantially reduce the likelihood of false annotations and provide users with a clear indication of annotation confidence. B.2AN ALTERNATIVE ATTEMPT FOR PROTEIN STRUCTURE INPUTS In our study, we used protein structural data predicted by AlphaFold2, the most accurate and widely adopted source of protein structural information, to provide structure inputs for our structure- 17 Published as a conference paper at ICLR 2026 Table 9: Partial evaluation on protein structure input from AlphaFold2 and Boltz-2. Subcellular Locations CDConv t (1,223 protein structures from AlphaFold2) CDConv t (1,223 protein structures from Boltz-2) PrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.7980.7090.7510.7880.5650.658 Nuclear Membrane------ Nucleoli0.5000.0640.1130.2670.0250.047 Nucleoplasm0.7940.6490.7140.7910.4480.572 Cytoplasm0.7010.5770.6330.6340.6730.653 Cytosol0.6450.4290.5150.5410.6220.579 Cytoskeleton0.5480.0850.1470.4670.0350.065 Centrosome------ Mitochondria0.6940.2000.3110.3580.1520.213 Endoplasmic Reticulum0.6300.1680.2660.5000.1490.229 Golgi Apparatus0.5000.0230.0440.3330.0080.015 Cell Membrane0.8230.3820.5220.8260.2690.406 Endosome------ Lipid Droplet------ Lysosome/Vacuole------ Peroxisome------ Vesicle0.8000.0190.0370.6670.0100.019 Primary Cilium------ Secreted Proteins0.7240.3750.4940.6550.3390.447 Sperm------ Micro Avg0.7450.4070.5270.6680.3710.477 Macro Avg0.4080.1840.2270.3410.1650.195 t The original MLP is replaced by Transformer layers. “–” indicates that no prediction is made for that location. based models. We acknowledge that Boltz is an efficient implementation of the still-unreleased AlphaFold3, and therefore seek to examine the subcellular localization performance when using Boltz-predicted structures as input (Wohlwend et al., 2025; Passaro et al., 2025). To the best of our knowledge, Boltz has not publicly released a complete set of inference results of human protein structures, compared with the AlphaFold2 dataset whose complete inference results can be downloaded publicly. Therefore, we locate a partial dataset of Boltz inference results, which includes structures of protein complexes predicted by Boltz-2 (Ille et al., 2025). From this work, we extracted the subset overlapping with our CAPSUL benchmark, yielding 1,223 proteins. For these 1,223 proteins, we compared different structure inputs (i.e., structures predicted by AlphaFold2 and by Boltz-2) on our previously trained structure-based model CDConv for inference. We report the results in Table 9. Across the overall metric and the majority of subcellular locations, AlphaFold2-based structural inputs outperform Boltz-based inputs. This observation further supports the rationale behind our use of AlphaFold2-predicted structures in constructing CAPSUL. As the most accurate and widely adopted source of protein structural information, AlphaFold2 provides high-quality structural inputs that lead to strong downstream performance in subcellular localization. CEXPERIMENT DETAILS C.1IMPLEMENTATION DETAILS The experiments were performed utilizing NVIDIA RTX 3090, A40 and A100 GPUs. We employ an early stopping strategy to mitigate overfitting with a tolerance of 5 epochs. Hyperparameters such as learning rate, number of epochs, and batch size are explored separately for each model type, considering their distinct architectures. C.2FURTHER DESCRIPTION OF STRUCTURE-BASED MODELS AND THEIR EXTENSIONS C.2.1OVERVIEW OF GRAPH CONSTRUCTION In the graphs constructed by our structure-based models, each node represents an amino acid, and edges encode the relationships between amino acids. Specifically, the edges are constructed as follows: 18 Published as a conference paper at ICLR 2026 Edge Criteria. There are two types of adjacency to form the edges in the graphs. Sequential adjacency refers to the proximity of amino acids along the one-dimensional primary sequence of a protein (e.g., if the sequential adjacency range is set to 3, then the amino acids from position [x-3, x+3] are considered sequential neighbors of the x-th residue). On the other hand, spatial adjacency captures the proximity of amino acids in the three-dimensional space of the protein (e.g., if the spatial adjacency radius is set to 8 ̊ A, all amino acids located within an 8 ̊ Asphere centered at a given residue are considered its spatial neighbors). These adjacency relationships define the edges in the constructed protein graph. Edge Features. For the GCN-based baselines CDConv and GearNet-Edge, we adopted their in- novative edge feature implementation methods originally proposed in their respective frameworks, which can be found in the corresponding publications (Fan et al., 2022; Zhang et al., 2022). These methods incorporate and encode both relative orientation and Euclidean distance. For our extended models, Graph Transformer, Graph Mamba, and Graph Diffusion, the edge features are derived by processing the aforementioned different edge criteria through an embedding layer. C.2.2EXPLANATION OF GRAPH ENCODER Within the structure-based baseline models, the graph encoders vary in their approaches to process- ing the input feature vectors: A GCN updates node representations via neighborhood aggregation, i.e., m (0) i = x i , m (l+1) i = σ P j∈N(i) W (l) m (l) j + b (l) . m (l) i is the representation of node i at layer l (the first layer is initialized with node embeddings x i various in different baselines), N (i) denotes the neighbors of node i, W (l) and b (l) are trainable weights and bias, and σ is a non-linear activation function (e.g., ReLU). After L layers of graph convolution, we obtain the final node repre- sentationsm (L) i n i=1 . To enhance interpretability and capture global interactions among residues, we replace the traditional average pooling with a Transformer encoder T (·) to obtain the residue representation, i.e., h = (h 1 ,..., h n ) =T m (L) i n i=1 , where h∈R n×d . Similarly, Graph Transformer and Graph Mamba substitute the convolution-based encoder with their respective architectures, while adhering to the same overall procedure to obtain the global protein representation. Given a protein structure represented as a graph, we introduce a diffusion-based refinement pro- cess in which node coordinates or geometric features are gradually perturbed with Gaussian noise and then denoised through a learned reverse process. The diffusion module serves as an auxiliary representation-learning stage designed to enhance geometric feature extraction prior to the down- stream subcellular localization prediction. This allows the network to capture multi-scale spatial dependencies while remaining robust to structural noise. C.2.3EXPLANATION OF CONTRASTIVE LEARNING AND FUSION MODELS CDConv with Contrastive Loss. For each of the 20 subcellular compartments, we construct posi- tive pairs (i.e., pairing protein samples that localize to the same compartment), and positive–negative pairs (i.e., pairing one protein that localizes to the compartment with another that does not). On top of the original loss function, we incorporate a contrastive loss to encourage higher embedding simi- larity for positive pairs while enforcing lower similarity for positive–negative pairs. We provide a formal mathematical description of how the contrastive loss is incorporated. Let ̄ h denote the protein-level representations from model (i.e., average embedding obtained before MLP classifier). For each of the 20 independent classification tasks, we compute the cosine similarity between representations of positive–positive pairs and positive–negative pairs, and construct the contrastive loss accordingly. For class c∈1,..., 20, let • P c =i| y i,c = 1 denote the set of positive samples, • N c =i| y i,c = 0 denote the negative samples, • ̄ h i denotes the protein-level embedding of the sample i. The positive–positive similarity matrix is S (+,+) ij = cos(z i , z j ), i,j ∈P c . 19 Published as a conference paper at ICLR 2026 We encourage positive samples to be close to each other by minimizing L (+,+) c = 1− 1 |P c | 2 X i,j∈P c S (+,+) ij . The positive–negative similarity matrix is S (+,−) ij = cos(z i , z j ), i∈P c ,j ∈N c . We encourage positive and negative samples to be dissimilar by minimizing L (+,−) c = 1 |P c ||N c | X i∈P c X j∈N c S (+,−) ij . Thus, the contrastive loss for class c is L contrast c =L (+,+) c +L (+,−) c . Finally, the overall contrastive loss averaged over all classes is L contrast = 1 C C X c=1 L contrast c . ESM-C+CDConv Fusion Model. In the early fusion model, the structural representations pro- duced by CDConv (without the additional Transformer architecture introduced by our paper) for each amino acid are added to the initial protein embedding of ESM-C. The combined representation is then passed through the pretrained ESM-C Transformer for interaction, followed by mean pool- ing to obtain a protein-level representation for downstream classification. The late fusion setting is similar, except that the structural representations from CDConv are added to the final sequence representation produced by ESM-C before mean pooling for the downstream classification task. C.3HYPERPARAMETER SETTINGS For all the experiments, we choose the best hyperparameters according to the best micro F1-score on the test set. For the main experiment, the best hyperparameter setting for each model is as follows: 1) ESM-2 (650M), the MLP hidden layers are set to (512,256), and learning rate to 1 × 10 −4 . 2) ESM- C (600M), the MLP hidden layers are set to (512,256) (to (512) when finetuning), and learning rate to 5× 10 −4 . 3) FoldSeek, the embedding dimensions are set to 256, transformer layers to 2, transformer heads to 4, and learning rate to 1× 10 −4 . 4) CDConv, the kernel channels are set to 24, feature channels to (256,512), geometric adjacency to 4 ̊ Aand 8 ̊ A(gradually increase with the convolutional layers), sequential adjacency to 5, sequential kernel size to 5, transformer layers to 3, transformer heads to 2, and learning rate to 5× 10 −4 . 5) GearNet-Edge, the max sequence length is set to 3,000, spatial adjacency set to [5 ̊ A,10 ̊ A], KNN adjacency set to [5,10], sequential adjacency set to 2, convolution hidden dimensions to (512,512,512), transformer layers to 2, transformer heads to 2, and learning rate to 1 × 10 −5 . 6) Graph Transformer, the transformer layers are set to 10, geometric adjacency to 10 ̊ A, sequential adjacency to 5, node dimensions set to 256, positional embedding dimension set to 8, and learning rate set to 5× 10 −5 . 7) Graph Mamba, the Mamba layers are set to 5, geometric adjacency to 10 ̊ A, sequential adjacency to 5, node dimensions set to 256, and learning rate set to 1× 10 −4 . 8) Graph Diffusion, the node dimensions are set to 64, geometric adjacency to 10 ̊ A, sequential adjacency to 5, timesteps set to 200, variance of Gaussian noise set to 1× 10 −4 in the beginning and 0.02 in the end, weight for diffusion loss to 0.1 compared with classification loss, and learning rate to 5 × 10 −4 . 9) CDConv with Contrastive Loss, the weight for contrastive loss is set to 0.1 compared with classification loss, and other settings are the same as the originial CDConv. 10) ESM-C+CDConv Fusion Model, the learning rate is set to 5× 10 −4 for early fusion and to 1× 10 −4 for late fusion, and other settings are the same as the originial ESM-C and CDConv. For the reweighting strategy, we inherit the optimal hyperparameter settings for ESM-C (600M), CDConv, and GearNet-Edge mentioned above. The best reweighting scheme for each model is as 20 Published as a conference paper at ICLR 2026 follows: 1) ESM-C (600M), focal loss with α set to the weights of log-inverse frequency, and γ set to 1.0. 2) CDConv, focal loss with α set to the weights of log-inverse frequency, and γ set to 3.0. 3) GearNet-Edge, inverse frequency reweighting. For the single-label classification strategy, we inherit the optimal hyperparameter settings for ESM-C (600M), CDConv, and GearNet-Edge mentioned above. To address class imbalance, we undersam- ple the negative class to achieve a 1:3 positive-to-negative sample ratio for ESM-C (600M), and a 1:1 positive-to-negative sample ratio for CDConv. DDETAILED BASELINE RESULTS Detailed experimental results of main experiments, reweighting strategy, and single-label classifica- tion strategy are provided in Tables 10, 11, and 12, respectively. They include evaluation metrics of precision, recall, and F1-score. EABLATION STUDY Although it has been recognized in the biological community that many patterns of subcellular localization cannot be fully captured by simple sequence information, we aim to investigate the potential benefits of incorporating protein structural information as input for prediction. Therefore, we conduct an ablation study on two representative structure-based baselines to quantify the positive impact of 3D information incorporated. Specifically, to preserve the integrity of the model input, we performed preprocessing on the pro- tein structural data. For each protein, we obtained the boundary values of its 3D coordinates and uniformly sampled the Cα coordinates at random within these boundaries to generate new protein structures. The 1D sequence data were kept unchanged, while the randomly sampled structures were used as the 3D structural input. Using the same hyperparameter settings as in the main experiments, we conducted an ablation study, with the detailed results shown in Table 13. We observed a signif- icant performance drop in this setting, which further demonstrates the decisive role of accurate 3D structural input in enabling correct model predictions. FEXPLANATION AND ILLUSTRATIVE EXAMPLES OF EVIDENCE-LEVEL ANNOTATIONS F.1EXPLANATION OF EVIDENCE-LEVEL ANNOTATIONS The evidence-level annotations design was originally intended to allow researchers to flexibly se- lect annotations based on their specific use cases. For instance, when the task is rigorous and re- quires high precision such as Nucleolar retention motifs discovery, selecting annotations with high confidence (i.e., choosing the experimentally validated annotations only) is more appropriate. Con- versely, for large-scale protein localization prediction, using lower-confidence but more abundant annotations (i.e., choosing both the non-experimentally validated and non-experimentally validated annotations) enriches the training data and leads to better model performance. F.2ILLUSTRATIVE EXAMPLES OF THREE STRATEGIES TOWARDS NON-EXPERIMENTALLY VALIDATED ANNOTATIONS In our main experiments, all non-experimentally validated annotations were treated as positive sam- ples to enhance the models’ performance in high-throughput prediction settings. Here we present two illustrative examples of the flexible usages of evidence-level annotations: 1) weighting labels (i.e., treating non-experimentally validated annotations as positive samples, but assigning a weight of 0.7 to them relative to experimental ones, which reduces the weight of non-experimental data in influencing the model) and 2) filtering labels (i.e., treating non-experimentally validated annota- tions as negative samples, which restricts models learning to experimental data with high reliability). The results are compared with the original one in our paper in Table 14. 21 Published as a conference paper at ICLR 2026 Table 10: Detailed performance of sequence-based and structure-based methods on CAPSUL. Subcellular Locations DeepLoc 2.1ESM-2 650MESM-2 650M f PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.6750.0860.152---0.6330.5860.609 Nuclear Membrane///------ Nucleoli///------ Nucleoplasm///---0.5920.5350.562 Cytoplasm0.5100.1000.167---0.5980.1570.248 Cytosol///------ Cytoskeleton///---0.2000.0030.006 Centrosome///------ Mitochondria0.7990.0650.120---0.8500.1950.317 Endoplasmic Reticulum0.5810.0670.121------ Golgi Apparatus0.5940.0320.061------ Cell Membrane0.7400.0780.142---0.7220.4510.555 Endosome///------ Lipid Droplet///------ Lysosome/Vacuole0.1980.0840.118------ Peroxisome0.6670.0730.131------ Vesicle///------ Primary Cilium///------ Secreted Proteins0.7730.1090.191---0.7420.6860.713 Sperm///------ Micro Avg///---0.6470.2640.375 Macro Avg///---0.2170.1310.150 Subcellular Locations ESM-C 600MESM-C 600M f ESM-C 600M 0 PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.6940.6090.6490.7080.5970.6480.6260.4980.555 Nuclear Membrane--------- Nucleoli0.8000.0480.0911.0000.0200.0391.0000.0120.024 Nucleoplasm0.6790.5730.6210.6860.5700.6230.6200.4180.500 Cytoplasm0.6110.4770.5360.6140.4990.5510.5070.3850.438 Cytosol0.5410.3070.3920.5670.2860.3800.4560.1040.169 Cytoskeleton0.6810.1540.2510.6290.1230.2050.4710.0250.048 Centrosome1.0000.0070.014------ Mitochondria0.8650.4160.5620.9030.3890.5440.6670.0530.099 Endoplasmic Reticulum0.6870.2350.3510.6740.2210.3330.5000.0310.059 Golgi Apparatus0.9380.0520.0991.0000.0140.027--- Cell Membrane0.7770.5310.6310.7530.5680.6480.7860.2430.372 Endosome1.0000.0090.018------ Lipid Droplet--------- Lysosome/Vacuole--------- Peroxisome--------- Vesicle1.0000.0050.009---1.0000.0020.005 Primary Cilium0.6820.0930.1640.5560.0620.112--- Secreted Proteins0.9030.7610.8260.8770.7300.7970.6040.3380.433 Sperm0.5000.0280.0520.6670.0370.070--- Micro Avg0.6900.3860.4950.6930.3820.4920.5980.2360.338 Macro Avg0.6180.2150.2630.4820.2060.2490.3620.1060.135 Subcellular Locations FoldSeekCDConv t GearNet-Edge t PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.6160.3980.4840.6510.5920.6200.6190.4500.521 Nuclear Membrane--------- Nucleoli---0.5830.0840.1470.5310.0680.121 Nucleoplasm0.5910.3410.4330.6330.5410.5830.6130.4440.515 Cytoplasm0.5810.1020.1740.5800.4140.4830.4980.4910.495 Cytosol0.5000.0010.0030.4890.2770.3530.4170.3580.385 Cytoskeleton0.4800.0380.0700.6490.0750.1350.2960.1860.228 Centrosome------0.2280.0880.127 Mitochondria---0.7070.3590.4760.4700.2400.318 Endoplasmic Reticulum---0.4410.2180.2920.4750.1970.279 Golgi Apparatus---0.7330.0380.0730.2110.0140.026 Cell Membrane0.6260.2370.3430.7210.4610.5620.7080.4570.556 Endosome------0.3640.0370.067 Lipid Droplet--------- Lysosome/Vacuole------0.4290.0400.073 Peroxisome--------- Vesicle---0.6670.0140.0270.2700.0390.068 Primary Cilium------0.4670.0870.147 Secreted Proteins0.6000.2250.3280.7950.7410.7670.7220.6550.687 Sperm------0.7140.0460.086 Micro Avg0.6050.1560.2480.6320.3520.4520.5460.3370.417 Macro Avg0.2000.0670.0920.3820.1910.2260.4020.1950.235 22 Published as a conference paper at ICLR 2026 Table 10: (Continued) Detailed performance of sequence-based and structure-based methods on CAPSUL. Subcellular Locations Graph TransformerGraph MambaGraph Diffusion PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.6640.5430.5970.5620.5560.5590.6170.6310.624 Nuclear Membrane---0.0610.0260.037--- Nucleoli0.5540.1240.2030.4330.1040.1680.6670.0240.047 Nucleoplasm0.6420.4830.5520.5260.4810.5020.5900.5660.578 Cytoplasm0.5520.3050.3930.4760.3730.4180.5420.4690.503 Cytosol0.4570.1700.2480.4210.4310.4260.4400.2140.288 Cytoskeleton0.5380.0220.0420.2490.2960.2700.3830.0570.099 Centrosome---0.1280.3130.1811.0000.0070.014 Mitochondria0.6880.3630.4750.4070.2940.3410.6800.1950.303 Endoplasmic Reticulum0.5520.1110.1840.5290.0310.0590.5560.0870.150 Golgi Apparatus0.8570.0210.0410.1820.1880.185--- Cell Membrane0.7180.4420.5470.4170.7660.5400.7370.3730.496 Endosome---0.1250.0830.100--- Lipid Droplet--------- Lysosome/Vacuole--------- Peroxisome--------- Vesicle0.5260.0230.0440.3060.0860.1350.6670.0090.018 Primary Cilium0.5000.0060.0120.2050.0560.088--- Secreted Proteins0.7670.6520.7050.4260.8020.5570.7380.5390.623 Sperm1.0000.0090.0180.1160.1470.130--- Micro Avg0.6370.3020.4100.4140.4080.4110.5960.3290.424 Macro Avg0.4510.1640.2030.2790.2520.2350.3810.1590.187 Subcellular Locations CDConv with Contrastive LossESM-C+CDConv Early FusionESM-C+CDConv Late Fusion PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.6810.5240.5920.7120.5870.6430.7010.5980.645 Nuclear Membrane--------- Nucleoli0.5560.0800.1400.7390.0680.1250.2100.1200.153 Nucleoplasm0.6460.4870.5560.7040.5910.6430.6760.5670.617 Cytoplasm0.5900.4040.4800.6820.3410.4550.5720.4690.515 Cytosol0.5160.1650.2500.4870.0940.1570.4670.3060.370 Cytoskeleton0.4730.1640.2430.7730.0530.1000.4120.2200.287 Centrosome------0.2310.0200.037 Mitochondria0.7540.3400.4680.8720.4160.5630.8270.4200.557 Endoplasmic Reticulum0.5190.2770.3610.5950.3560.446--- Golgi Apparatus0.5530.0910.1560.4620.1710.249--- Cell Membrane0.7460.4220.5390.7740.5300.6290.7320.6220.673 Endosome0.2500.0190.0341.0000.0280.054--- Lipid Droplet--------- Lysosome/Vacuole---1.0000.0130.026--- Peroxisome--------- Vesicle0.7500.0140.0270.4320.0360.067--- Primary Cilium0.5000.0190.036---0.2550.0750.115 Secreted Proteins0.7970.7650.7800.9050.7470.8190.9080.6040.725 Sperm1.0000.0090.018------ Micro Avg0.6500.3260.4350.7100.3510.4700.6340.3810.476 Macro Avg0.4670.1890.2340.5070.2020.2490.3000.2010.235 f We finetune the pre-trained protein language model. t The original MLP is replaced by Transformer layers. 0 The parameters of ESM-C is initialized randomly. “/” indicates that DeepLoc 2.1 does not support prediction for that location, and therefore, average metrics are not considered in this case. “–” indicates that no prediction is made for that location. 23 Published as a conference paper at ICLR 2026 Table 11: Detailed performance of selected baselines with reweighting scheme. Subcellular Locations ESM-C 600MCDConv t GearNet-Edge t PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.6980.5750.6300.4810.8920.6250.4840.8560.618 Nuclear Membrane---0.0330.5660.0620.0460.0790.058 Nucleoli---0.1050.9160.1880.1530.4180.224 Nucleoplasm0.6790.5000.5760.4690.8590.6070.4360.8410.574 Cytoplasm0.5680.4460.5000.4500.8230.5820.4410.7110.544 Cytosol0.5130.0760.1330.3530.8290.4950.3660.7140.484 Cytoskeleton0.7780.0440.0830.1840.6980.2920.2180.4500.294 Centrosome---0.0890.7760.1600.1340.2520.175 Mitochondria0.8460.3360.4810.1910.6720.2970.2470.4270.313 Endoplasmic Reticulum ---0.1950.7370.3080.2760.4600.345 Golgi Apparatus---0.1520.6480.2460.1770.3660.238 Cell Membrane0.7230.4650.5660.4620.7090.5600.3980.8200.536 Endosome---0.0670.4070.1140.1770.1300.150 Lipid Droplet1.0000.1330.2350.0140.0670.0230.3330.0670.111 Lysosome/Vacuole---0.1170.3470.1750.1160.1070.111 Peroxisome1.0000.1050.1900.0400.4210.0720.1110.1050.108 Vesicle---0.1980.5320.2880.2060.4450.281 Primary Cilium0.6670.0120.0240.0960.6400.1670.1230.3110.176 Secreted Proteins0.8330.7300.7780.4130.8910.5640.5090.7750.614 Sperm---0.0660.6790.1200.1090.1470.125 Micro Avg0.6790.3130.4290.2530.7720.3810.3480.6500.453 Macro Avg0.4150.1710.2100.2090.6550.1970.2530.4240.304 t The original MLP is replaced by Transformer layers. “–” indicates that no prediction is made for that location. Table 12: Detailed performance of selected baselines with single-label classification strategy. Subcellular Locations ESM-C 600MCDConv t GearNet-Edge t PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nuclear Membrane---0.0270.7110.0520.0260.1180.042 Nucleoli0.2510.2850.2670.0820.9920.1510.1510.4700.228 Centrosome0.1240.3610.1840.0510.3330.0890.0990.5310.167 Golgi Apparatus0.2930.2680.2800.0800.1990.1140.1610.3030.210 Endosome0.1110.3330.1670.0290.1760.0490.0820.2780.126 Lipid Droplet0.0110.2000.021---0.0320.1330.051 Lysosome/Vacuole0.0750.2530.115---0.0970.4930.162 Peroxisome0.0290.5260.054---0.0130.1580.023 Vesicle0.2700.0390.0680.1410.6250.2300.2070.3800.268 Primary Cilium0.1750.4600.2530.0550.3790.0970.1040.4720.171 Sperm0.1210.2290.1590.0450.1380.0680.0770.2390.117 t The original MLP is replaced by Transformer layers. “–” indicates that no prediction is made for that location. Table 13: Detailed performance comparison of CDConv and GearNet-Edge under random sampling of Cα coordinates. Subcellular Locations CDConv t (ablation)CDConv t GearNet-Edge t (ablation)GearNet-Edge t PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.5950.5120.5500.6510.5920.6200.5150.4590.4850.6190.4500.521 Nuclear Membrane------------ Nucleoli0.4170.0200.0380.5830.0840.1470.2680.0760.1190.5310.0680.121 Nucleoplasm0.5780.4190.4860.6330.5410.5830.4790.4280.4520.6130.4440.515 Cytoplasm0.5350.2140.3060.5800.4140.4830.4320.4140.4220.4980.4910.495 Cytosol0.4780.0690.1200.4890.2770.3530.3940.2790.3270.4170.3580.385 Cytoskeleton---0.6490.0750.1350.2520.1190.1620.2960.1860.228 Centrosome------0.1840.0610.0920.2280.0880.127 Mitochondria0.5370.0840.1450.7070.3590.4760.2830.0650.1060.4700.2400.318 Endoplasmic Reticulum 0.6250.0170.0340.4410.2180.2920.3210.0620.1040.4750.1970.279 Golgi Apparatus---0.7330.0380.0730.1320.0170.0310.2110.0140.026 Cell Membrane0.6210.4250.5050.7210.4610.5620.5950.4130.4870.7080.4570.556 Endosome---------0.3640.0370.067 Lipid Droplet------------ Lysosome/Vacuole---------0.4290.0400.073 Peroxisome------------ Vesicle---0.6670.0140.0270.1850.0770.1090.2700.0390.068 Primary Cilium------0.3680.0430.0780.4670.0870.147 Secreted Proteins0.7030.2180.3330.7950.7410.7670.5150.2320.3200.7220.6550.687 Sperm------0.1250.0090.0170.7140.0460.086 Micro Avg0.5860.2290.3290.6320.3520.4520.4500.2830.3480.5460.3370.417 Macro Avg0.2540.0990.1260.3820.1910.2260.2520.1380.1660.4020.1950.235 t The original MLP is replaced by Transformer layers. “–” indicates that no prediction is made for that location. 24 Published as a conference paper at ICLR 2026 Table 14: Detailed performance of two illustrative examples of evidence-level annotations: weight- ing labels and filtering labels. Subcellular Locations ESM-C 600MESM-C 600M (weighting)ESM-C 600M (filtering) PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.6940.6090.6490.7170.5850.6450.7030.5750.633 Nuclear Membrane--------- Nucleoli0.8000.0480.0910.7860.0880.1590.8670.0550.104 Nucleoplasm0.6790.5730.6210.6920.5580.6180.6860.5510.611 Cytoplasm0.6110.4770.5360.6190.4530.5230.5720.4000.470 Cytosol0.5410.3070.3920.5340.2560.3460.5900.2390.340 Cytoskeleton0.6810.1540.2510.6490.1160.1970.3810.0340.062 Centrosome1.0000.0070.0141.0000.0070.014--- Mitochondria0.8650.4160.5620.9070.3740.5300.7980.3730.508 Endoplasmic Reticulum0.6870.2350.3510.7260.1560.2560.5000.0390.072 Golgi Apparatus0.9380.0520.0991.0000.0100.021--- Cell Membrane0.7770.5310.6310.7570.5700.6500.6610.2540.367 Endosome1.0000.0090.018------ Lipid Droplet--------- Lysosome/Vacuole--------- Peroxisome--------- Vesicle1.0000.0050.0091.0000.0050.009--- Primary Cilium0.6820.0930.1640.5380.0430.0801.0000.0080.016 Secreted Proteins0.9030.7610.8260.9200.6690.775--- Sperm0.5000.0280.0520.8000.0370.0700.5000.0100.019 Micro Avg0.6900.3860.4950.7000.3660.4810.6570.3060.418 Macro Avg0.6180.2150.2630.5820.1960.2450.3630.1270.160 Subcellular Locations CDConv t CDConv t (weighting)CDConv t (filtering) PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.6510.5920.6200.6740.5390.5990.6440.5300.581 Nuclear Membrane--------- Nucleoli0.5830.0840.1470.5560.0200.0390.4840.6400.112 Nucleoplasm0.6330.5410.5830.6570.4970.5660.6520.4240.514 Cytoplasm0.5800.4140.4830.5570.4690.5090.5280.4450.483 Cytosol0.4890.2770.3530.4700.2930.3610.4680.3370.392 Cytoskeleton0.6490.0750.1350.7140.0630.1160.4000.0080.016 Centrosome--------- Mitochondria0.7070.3590.4760.7620.3550.4840.5880.3630.449 Endoplasmic Reticulum0.4410.2180.2920.5610.1590.2480.2780.0240.045 Golgi Apparatus0.7330.0380.0730.6150.0280.053--- Cell Membrane0.7210.4610.5620.7230.4470.5530.5620.1940.288 Endosome--------- Lipid Droplet--------- Lysosome/Vacuole--------- Peroxisome--------- Vesicle0.6670.0140.027------ Primary Cilium---0.3330.0060.012--- Secreted Proteins0.7950.7410.7670.8570.5730.6870.4000.0440.079 Sperm--------- Micro Avg0.6320.3520.4520.6370.3330.4380.5770.2910.386 Macro Avg0.3820.1910.2260.3740.1730.2110.2500.1220.148 Subcellular Locations GearNet-Edge t GearNet-Edge t (weighting)GearNet-Edge t (filtering) PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.6190.4500.5210.6220.3370.4370.6070.4780.535 Nuclear Membrane--------- Nucleoli0.5310.0680.1210.4210.0320.0600.3330.0640.107 Nucleoplasm0.6130.4440.5150.6180.3260.4270.5760.4530.507 Cytoplasm0.4980.4910.4950.4730.4660.4690.4520.4790.465 Cytosol0.4170.3580.3850.4240.3390.3770.3940.4030.398 Cytoskeleton0.2960.1860.2280.3420.1670.2240.2840.1050.153 Centrosome0.2280.0880.1270.2130.0680.1030.2000.0540.085 Mitochondria0.4700.2400.3180.6670.1760.2780.5080.1420.221 Endoplasmic Reticulum0.4750.1970.2790.4350.1280.1980.2680.0730.115 Golgi Apparatus0.2110.0140.0260.2110.0140.0260.2500.0040.008 Cell Membrane0.7080.4570.5560.6290.3920.4830.4630.1700.248 Endosome0.3640.0370.0670.3330.0280.0510.5000.0290.056 Lipid Droplet--------- Lysosome/Vacuole0.4290.0400.0730.1250.0130.024--- Peroxisome--------- Vesicle0.2700.0390.0680.2680.0930.1380.2800.0550.092 Primary Cilium0.4670.0870.1470.5240.0680.1210.3640.0310.058 Secreted Proteins0.7220.6550.6870.8260.5190.6370.2730.0660.106 Sperm0.7140.0460.0860.6670.0370.0700.5000.0200.038 Micro Avg0.5460.3370.4170.5290.2820.3680.4850.3000.371 Macro Avg0.4020.1950.2350.3900.1600.2060.3130.1310.160 t The original MLP is replaced by Transformer layers. “–” indicates that no prediction is made for that location. 25 Published as a conference paper at ICLR 2026 Table 15: Detailed performance of different weights for non-experimentally validated annotations on ESM-C. Subcellular Locations ESM-C weighting 0.1ESM-C weighting 0.3ESM-C weighting 0.5 PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.7330.5490.6280.730.5530.6290.7280.5640.636 Nuclear Membrane--------- Nucleoli0.7890.0600.1120.7920.0760.1390.8240.0560.105 Nucleoplasm0.7080.5180.5980.7140.5190.6010.7100.5220.602 Cytoplasm0.6050.5110.5540.5920.5060.5460.6000.5050.548 Cytosol0.5300.3500.4220.5180.3360.4080.5110.3820.437 Cytoskeleton0.6360.1100.1880.5870.1160.1940.6180.1070.182 Centrosome------1.0000.0070.014 Mitochondria0.9170.3820.5390.9170.3820.5390.9140.3660.523 Endoplasmic Reticulum0.7500.1140.1980.6980.1280.2160.6960.1660.268 Golgi Apparatus------1.0000.0030.007 Cell Membrane0.7710.3670.4970.7680.3670.4960.7910.4420.567 Endosome--------- Lipid Droplet--------- Lysosome/Vacuole--------- Peroxisome--------- Vesicle------1.0000.0050.009 Primary Cilium0.5830.0430.0810.6470.0680.1240.6250.0310.059 Secreted Proteins0.9330.4740.6290.9360.4470.6050.9580.4710.632 Sperm0.1670.0090.0170.3750.0280.0510.6670.0180.036 Micro Avg0.6870.3380.4530.6820.3380.4520.6850.3530.466 Macro Avg0.4060.1740.2230.4140.1760.2270.5820.1820.231 Subcellular Locations ESM-C weighting 0.7ESM-C weighting 0.9 PrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.7170.5850.6450.7110.5780.638 Nuclear Membrane------ Nucleoli0.7860.0880.1590.8330.0800.147 Nucleoplasm0.6920.5580.6180.6990.5480.614 Cytoplasm0.6190.4530.5230.5960.5170.554 Cytosol0.5340.2560.3460.5170.3720.432 Cytoskeleton0.6490.1160.1970.6460.1320.219 Centrosome1.0000.0070.0141.0000.0070.014 Mitochondria0.9070.3740.5300.8990.3740.528 Endoplasmic Reticulum0.7260.1560.2560.6800.2350.350 Golgi Apparatus1.0000.0100.0210.9090.0350.067 Cell Membrane0.7570.5700.6500.7590.5400.631 Endosome---0.6670.0190.036 Lipid Droplet------ Lysosome/Vacuole------ Peroxisome------ Vesicle1.0000.0050.0091.0000.0070.014 Primary Cilium0.5380.0430.0800.5710.0500.091 Secreted Proteins0.9200.6690.7750.9070.7340.811 Sperm0.8000.0370.0700.7500.0280.053 Micro Avg0.7000.3660.4810.6830.3880.495 Macro Avg0.5820.1960.2450.6070.2130.260 “–” indicates that no prediction is made for that location. As shown in the table, models that treat non-experimentally validated annotations as positive sam- ples generally achieve the best overall performance. This may be because many non-validated anno- tations also originate from reliable sources (e.g., the ”ECO:0000303” code in the UniProt database indicates that the localization information is extracted from published literature); thus, they still hold relatively high credibility. This also demonstrates that for large-scale deep learning training, as our work does, including such annotations can increase sample diversity and improve data richness, and therefore, improve models’ performance. F.3ANALYSIS OF DIFFERENT WEIGHTS FOR NON-EXPERIMENTALLY VALIDATED ANNOTATIONS To have a comprehensive analysis of the effect of evidence level, we further extend the weighting- label strategy by setting different weights to samples with non-experimentally validated annotations. We report the results in Table 15 and 16. From the results, we have the following observation: 26 Published as a conference paper at ICLR 2026 Table 16: Detailed performance of different weights for non-experimentally validated annotations on CDConv. Subcellular Locations CDConv t weighting 0.1CDConv t weighting 0.3CDConv t weighting 0.5 PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.6400.6460.6430.6160.6620.6380.6710.5430.600 Nuclear Membrane--------- Nucleoli0.6520.0600.1100.4710.0960.1600.6670.0160.031 Nucleoplasm0.6140.5330.5710.5960.6000.5980.6400.5090.567 Cytoplasm0.5940.4500.5120.5920.3980.4760.5580.5140.535 Cytosol0.4930.3220.3900.4780.2780.3520.4680.3480.399 Cytoskeleton0.7500.0090.0190.8180.0280.0550.6190.0410.077 Centrosome--------- Mitochondria0.7280.3470.4700.7080.3700.4860.7030.3440.462 Endoplasmic Reticulum0.5590.1310.2130.4750.1310.2060.5400.1630.250 Golgi Apparatus------0.7500.0100.021 Cell Membrane0.7620.3330.4630.7320.4130.5280.7530.4450.560 Endosome--------- Lipid Droplet--------- Lysosome/Vacuole--------- Peroxisome--------- Vesicle--------- Primary Cilium1.0000.0060.0121.0000.0060.012--- Secreted Proteins0.8230.6180.7060.8350.6720.7450.8680.5390.665 Sperm--------- Micro Avg0.6310.3400.4420.6170.3540.4500.6290.3430.444 Macro Avg0.3810.1730.2050.3660.1830.2130.3620.1740.208 Subcellular Locations CDConv t weighting 0.7CDConv t weighting 0.9 PrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.6740.5390.5990.6820.5230.592 Nuclear Membrane------ Nucleoli0.5560.0200.0390.7500.0120.024 Nucleoplasm0.6570.4970.5660.6690.4790.558 Cytoplasm0.5570.4690.5090.5510.5710.561 Cytosol0.4700.2930.3610.4700.3980.431 Cytoskeleton0.7140.0630.1160.4800.0750.130 Centrosome------ Mitochondria0.7620.3550.4840.7070.3320.452 Endoplasmic Reticulum0.5610.1590.2480.4810.1730.254 Golgi Apparatus0.6150.0280.0530.5880.0350.066 Cell Membrane0.7230.4470.5530.7250.4930.587 Endosome------ Lipid Droplet------ Lysosome/Vacuole------ Peroxisome------ Vesicle------ Primary Cilium0.3330.0060.0120.3330.0060.012 Secreted Proteins0.8570.5730.6870.8450.5970.700 Sperm------ Micro Avg0.6370.3330.4380.6250.3590.456 Macro Avg0.3740.1730.2110.3640.1850.218 t The original MLP is replaced by Transformer layers. “–” indicates that no prediction is made for that location. Table 17: Analysis of precision and recall on Nucleus on ESM-C. ESM-C weighting0.10.30.50.70.9Treat as Positive Precision for Nucleus0.7330.7300.7280.7170.7110.694 Recall for Nucleus0.5490.5530.5640.5850.5780.609 1) For certain subcellular compartments, baselines trained exclusively on experimentally validated annotations or assigned low positive weights to non-experimentally validated annotations generally achieve higher precision, as shown in Table 17. Actually, precision and recall often represent a trade-off in modeling strategies. That is, adopting a more conservative prediction strategy typically increases precision but reduces the number of correctly recalled samples, and vice versa. Therefore, selecting high-confidence evidence levels can be seen as a method of enforcing a more conser- vative prediction approach, helping to reduce the likelihood of false-positive predictions. This highlights the novelty of evidence-level annotations: using experimentally validated data is con- sidered a strategy to ensure high precision and reliability. 2) However, for overall results and most subcellular compartments, the results among the three strategies show no significant differences. This indirectly supports the notion we mentioned above that even non-validated annotations in CAPSUL still possess relatively high reliability. This 27 Published as a conference paper at ICLR 2026 Table 18: Detailed performance of hierarchical classifiers on CDConv. CDConv t Hierarchical Classifiers Subcellular LocationsPrecisionRecallF1-ScoreSubcellular LocationsPrecisionRecallF1-Score Nucleus’s classifierCytoplasm’s classifier Nuclear Membrane---Cytosol0.2600.9800.411 Nucleoli---Cytoskeleton0.3310.1260.182 Nucleoplasm0.3391.0000.507 t The original MLP is replaced by Transformer layers. “–” indicates that no prediction is made for that location. Table 19: Analysis of performance of three pLDDT groups on CDConv. Subcellular Locations High pLDDT GroupMedium pLDDT GroupLow pLDDT Group PrecisionRecallF1-ScorePrecisionRecallF1-ScorePrecisionRecallF1-Score Nucleus0.5610.3930.4620.6360.5700.6010.7000.7420.720 Nuclear Membrane--------- Nucleoli0.6670.0820.1460.6320.1360.2240.3750.0340.062 Nucleoplasm0.5120.3100.3870.5750.4970.5330.7130.7170.715 Cytoplasm0.6010.5130.5530.5780.4220.4880.5460.2980.385 Cytosol0.5200.4510.4830.4670.2410.3180.4140.1120.177 Cytoskeleton0.8330.0600.1120.5770.1280.2100.8000.0340.065 Centrosome--------- Mitochondria0.7890.4750.5930.6440.3220.4300.5290.1670.254 Endoplasmic Reticulum 0.5060.3440.4090.4100.1520.2220.1760.0540.082 Golgi Apparatus0.8570.0570.1070.7140.0480.089--- Cell Membrane0.8060.4630.5880.7430.5810.6530.5300.2720.360 Endosome--------- Lipid Droplet--------- Lysosome/Vacuole--------- Peroxisome---0.3330.0060.012--- Vesicle0.8330.0350.068------ Primary Cilium---0.8560.7200.782--- Secreted Proteins0.7460.7750.760---0.8160.7020.755 Sperm------ Micro Avg0.6170.3450.4430.6280.3450.4460.6510.3670.469 Macro Avg0.4120.1980.2330.5380.1910.2280.2800.1570.179 “–” indicates that no prediction is made for that location. ensures that, when CAPSUL is used for downstream tasks, the inclusion of non-experimentally validated annotations does not introduce substantial bias. GATTEMPT TOWARD HIERARCHICAL CLASSIFIERS FOR NESTED CATEGORIES In order to make precise predictions of nested subcategories, we attempt to train a hierarchical classifier for the two nested categories, Nucleus and Cytoplasm. These classifiers were built upon CDConv, using positive samples of the corresponding parent categories as the training set. We report the experimental results in Table 18. We have the following observations: 1) For the three subcategories within the nucleus, the hierarchical classifier actually leads to de- creased performance compared with the original CDConv. We attribute this to the reduced number of training samples when training this hierarchical classifier, which, combined with the already se- vere class imbalance in these subcategories, likely exacerbated the issue. Considering the results discussed in Section 4.3 of the main text, we find that for such highly imbalanced subcellular local- ization tasks, directly applying a single-label classification strategy, while adjusting the correspond- ing reweighting and classification thresholds, yields more significant improvements. 2) For the two subcategories within the cytoplasm, classification performance greatly improves, demonstrating that the hierarchical classification strategy can be beneficial for certain subcellular compartments. This result provides an additional strategy, hierarchical classifiers, to address the minority class issue apart from what we have discussed in Section 4.3 of the main text. 28 Published as a conference paper at ICLR 2026 HANALYSIS OF THE PLDDT SCORES OF PROTEIN STRUCTURE INPUT In AlphaFold-predicted structures, regions with low pLDDT often indicate intrinsically disordered segments, which naturally lack stable tertiary structure and are therefore harder for any predictor to model with high confidence. We conduct an analysis on the representative structure-based model, CDConv, examining the relationship between residue-level pLDDT and model performance across the test set. Specifically, we divide all 3,028 proteins in the test set into three groups of equal size based on their protein-level mean pLDDT values. The highest and lowest mean pLDDT values within each group are [83.56, 99.39], [70.40, 83.54], and [28.11, 70.40]. We then compare the performance of these groups to determine whether substantial performance discrepancies exist, which would suggest that the model is sensitive to variations in structural confidence. We report the results of three groups in Table 19. Upon further examination of specific subcellular categories, we find that: 1) For some subcellular locations (e.g., Nucleus, Cytoplasm, Cytosol, and Cell Membrane), the performance differences among the three groups may appear larger. This is probably attributable to the uneven distribution of positive samples across the groups. For example, in the Nucleus category, the high-pLDDT group contains only 318 positive samples, whereas the low-pLDDT group contains 476 positive samples, which contributes to this noticeable discrepancy. 2) For other subcellular locations (e.g., Cytoskeleton and Secreted Proteins), the results across the three groups do not exhibit pronounced differences. This observation is consistent with the trend shown by the overall evaluation metrics above, indicating that proteins with different average pLDDT levels have not greatly influenced the prediction of these subcellular localizations. In summary, the low-pLDDT group does not exhibit worse predictions systematically. This in- dicates that the CAPSUL is robust to fluctuations in pLDDT and does not rely disproportionately on regions of high structural confidence. In other words, CAPSUL provides stable and reliable struc- tural inputs for downstream tasks even when proteins contain disordered or low-confidence regions, demonstrating the quality and suitability of our structural dataset for subcellular lo- calization prediction. However, results of particular subcellular compartments highlight the need for further exploration of developing more informative and robust structural representations to better capture the determinants of subcellular localization in future work. ISAMPLE EFFICIENCY A sample efficiency curve would explicitly reflect a possible strategy to choose the proper size of the training set, which strikes a balance between model performance and computational costs. We randomly select training subsets of varying sizes and evaluate their sample efficiency curve on CDConv in Figure 3. We observed that the smaller dataset, compared with the original CAPSUL, exhibits noticeably poorer performance, reflected in both lower micro F1-scores and fewer categories for which any positive samples were successfully predicted. This indicates that, in order to achieve satisfactory overall performance as well as decent performance on minority classes, the current size of training samples is relatively appropriate. JINTERPRETABILITY WITH ATTENTION SCORE In Transformer architectures, the attention mechanism allows each token to compute a weighted representation of all other tokens in the sequence. Specifically, for a given token, a set of attention weights is derived via scaled dot-product operations between its query vector and the key vectors of all tokens, followed by a softmax normalization. These attention weights reflect how much in- formation the token attends to from each of its peers. To assess the relative importance of each token within the sequence, we aggregated the attention it receives from all other tokens, i.e., sum- ming over the attention scores directed toward that token across the entire sequence. This provides a global measure of how influential a token is in shaping the contextual representations learned by 29 Published as a conference paper at ICLR 2026 Figure 3: Sample efficiency curve on CDConv. Figure 4: Visualization of full attention scores and structures of proteins MFNG, B3GALT2, and GIMAP1, where the residues of known pattern α-helix are highlighted. the model. We interpret this aggregated attention as a proxy for biological interpretability, where highly attended residues may correspond to structurally or functionally important positions within the protein. In Section 4.3.3 of the main text, we introduce a CDConv model for predicting Golgi apparatus localization. By analyzing the attention score within the model’s Transformer architecture, we iden- tify a localization pattern associated with an α-helix, which is consistent with existing biological findings. Here, we visualize the full attention score of the three example proteins discussed in the main text (i.e., MFNG, B3GALT2, and GIMAP1), as shown in Figure 4. The residues of known lo- calization patterns α-helix are highlighted in orange for clear comparison. Notably, the 20 residues with the highest attention scores exhibit a 90% overlap with the ground truth, further highlighting the CDConv model’s precision in identifying localization patterns. 30 Published as a conference paper at ICLR 2026 More interpretability results. There are two more potential localization patterns identified on CAPSUL: 1) ”W-pair” for Golgi apparatus (i.e., two spatially adjacent Tryptophan residues), which aligns with existing studies suggesting that Tryptophan can influence Golgi targeting (Ashlin et al., 2021), and 2) a flexible region at the N-terminus of the protein for Mitochondria, which aligns with existing studies suggesting that this area can influence multiple compartments target- ing (Sohn et al., 2009). However, these newly identified patterns still require further experimental validation. Nevertheless, these findings together demonstrate that both our curated dataset and the selected structure-based baseline model are capable of capturing meaningful subcellular localization signals, offering promising insights and directions for future research in cell biology. KGENERALIZATION ABILITY The CAPSUL workflow is readily applicable to other species. While our current study exclu- sively uses human protein data, we do not assert that CAPSUL’s pipeline is intrinsically limited to human proteins. Our dataset and benchmark construction pipeline, including 1) the acquisition of one-dimensional and three-dimensional protein information, 2) the collection of localization anno- tations, and 3) the comparison of various baselines, is entirely species-agnostic. Once the CAPSUL pipeline becomes robust and well-received, there can be attempts to extend the same workflow to develop subcellular localization datasets and benchmarks across multiple species in parallel. This extension will not only enable broader biological investigations but also provide valuable resources for studying the evolutionary conservation and divergence of protein localization mechanisms across phylogenetic lineages. By comparing subcellular patterns across species, researchers can gain in- sights into the selective pressures shaping cellular organization, the emergence of organelle-specific functions, and the molecular adaptations that underpin evolutionary innovation. The CAPSUL workflow is applicable to context-dependent localization if sufficient related data is accessible. With regard to the context-dependent dynamic localization, current protein databases lack extensive data on protein structures across diverse biological contexts and their cor- responding variations in subcellular localization. If CAPSUL were augmented with dynamic data capturing protein structural changes (e.g., conformational shifts induced by stressors or ligands) from updated protein databases in the future, the baselines can be retrained on CAPSUL to model context-dependent localization. This would enable the models to make precise predictions about subcellular localization under specific cellular conditions or in response to external environmental states, moving beyond static localization to capture functional biological dynamics. The CAPSUL is readily extensible to additional protein data and baseline methods in the fu- ture. Our highly standardized pipeline ensures that any newly released protein data can be updated to CAPSUL. Also, CAPSUL embraces a broader range of other possible protein structural inputs (e.g., which can be used in the construction of graph nodes and edges) to enrich both the local and global representations of proteins. Moreover, the results of more advanced protein representation methods can be included to perform this downstream task if their input data is supported by CAP- SUL, which enables fair and informative comparison across various baseline methods. LAVAILABILITY OF DATASET AND CODE The complete dataset, including localization labels, extracted protein structures, etc. can be ac- cessed at https://huggingface.co/datasets/getbetterhyccc/CAPSUL. Our im- plementation is publicly available at https://github.com/getbetter-hyccc/CAPSUL. For some baseline models, we adopt publicly released implementations, including Graph Trans- former at https://github.com/pyg-team/pytorch_geometric/tree/master and Graph Mamba at https://github.com/alxndrTL/mamba.py. MTHE USE OF LARGE LANGUAGE MODELS In this study, Large Language Models (LLMs) were employed solely for linguistic refinement, such as polishing the clarity, grammar, and fluency of the manuscript. Importantly, all conceptual ad- vances, methodological innovations, experimental designs, and primary contributions presented in this work were independently conceived, developed, and validated by the authors. The role of LLMs 31 Published as a conference paper at ICLR 2026 was thus limited to improving readability and ensuring the precision of academic writing, without influencing the scientific content or originality of the research. 32