Paper deep dive
PhenSPINE: A Standardized Benchmark for Spine Pathology Diagnosis
Duong Ngoc Vu, Hai Son Nguyen, Trong-Nghia Nguyen, Bien Tran Van, Trang Mai Xuan, Huan Vu, Thien Van Luong
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/23/2026, 2:32:59 AM
Summary
The paper introduces PhenSPINE, a standardized benchmark dataset for spine pathology diagnosis comprising 16,813 MRI images from 250 patients. It proposes a deep learning benchmark integrating convolutional backbones with Positional Encoding to model intervertebral disc anatomy. Experiments across four MRI sequences reveal that the Sagittal T2-weighted (SAG-T2) sequence with EfficientNet-B1 achieves the best diagnostic performance (Macro F1-score 50.31%), while multi-sequence fusion strategies underperform due to noise interference.
Entities (15)
Relation Signals (12)
PhenSPINE → contains → Disc Herniation
confidence 95% · The dataset comprises ... Disc Herniation is observed in 6.92% of the dataset.
PhenSPINE → contains → Spondylolisthesis
confidence 95% · Spondylolisthesis ... appearing in only 3.29% ... of records
PhenSPINE → contains → Disc Narrowing
confidence 95% · Disc Narrowing are relatively rare, appearing in only 4.73% of records
PhenSPINE → contains → Disc Bulging
confidence 95% · Disc Bulging is the most prevalent condition, occurring in 26.75% of cases.
EfficientNet-B1 → achievesbestperformancewith → Sagittal T2-weighted
confidence 92% · the EfficientNet-B1 backbone, when paired with the SAG-T2 sequence, achieves the highest Macro F1-score of 50.31%
Positional Encoding → improves → EfficientNet-B1
confidence 90% · integration of Positional Encoding elevates Precision to 60.01% and the F1-score to 50.31%
Sagittal T2-weighted → outperforms → AX-T2
confidence 90% · SAG-T2 ... substantially outperforming ... AX-T2
Sagittal T2-weighted → outperforms → SAG-T1
confidence 90% · SAG-T2 ... substantially outperforming ... SAG-T1
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The accurate diagnosis of spinal pathologies depends heavily on radiological interpretation, yet automated systems are hindered by the lack of diverse, high-quality benchmarks. In this study, we present PhenSPINE, a Magnetic Resonance Imaging dataset comprising 16,813 images from 250 patients, curated to facilitate advanced deep learning research. We propose a robust diagnostic benchmark that integrates state-of-theart convolutional backbones with a Positional Encoding mechanism to explicitly model the anatomical context of intervertebral discs. Evaluating across four standard MRI sequences, our experiments demonstrate that the Sagittal T2-weighted sequence offers the most robust diagnostic value, achieving a superior Macro F1-score of 50.31%. We find that multisequence fusion strategies yield inferior performance compared to this single-sequence baseline, as the images across sequences in our dataset are significantly compromised by noise interference from surrounding anatomical regions. This work establishes a robust baseline and offers critical insights into sequence selection for spine analysis.
Tags
Links
- Source: https://arxiv.org/abs/2607.19696v1
- Canonical: https://arxiv.org/abs/2607.19696v1
Trouble viewing inline? Open PDF directly →
Full Text
28,603 characters extracted from source content.
Expand or collapse full text
11institutetext: Business AI Lab, College of Technology, National Economics University, Vietnam 22institutetext: A2I Lab, Phenikaa School of Computing, Phenikaa University, Hanoi, Vietnam 33institutetext: Medical Imaging & Radiological Technology Department, Faculty of Medical Technology, Phenikaa School of Medicine & Pharmacy, Phenikaa University, Vietnam 44institutetext: Radiology and Functional Exploration Center, Phenikaa University Hospital, Vietnam 44email: duongvn.bai@st.neu.edu.vn, hains24206@gmail.com, nghiant@neu.edu.vn, bien.tranvan@phenikaa-uni.edu.vn, trang.maixuan@phenikaa-uni.edu.vn, huanv, thienlv@neu.edu.vn PhenSPINE: A Standardized Benchmark for Spine Pathology Diagnosis Duong Ngoc Vu Hai Son Nguyen Trong-Nghia Nguyen Bien Tran Van Trang Mai Xuan Corresponding author Huan Vu Thien Van Luong Abstract The accurate diagnosis of spinal pathologies depends heavily on radiological interpretation, yet automated systems are hindered by the lack of diverse, high-quality benchmarks. In this study, we present PhenSPINE, a Magnetic Resonance Imaging dataset comprising 16,813 images from 250 patients, curated to facilitate advanced deep learning research. We propose a robust diagnostic benchmark that integrates state-of-the-art convolutional backbones with a Positional Encoding mechanism to explicitly model the anatomical context of intervertebral discs. Evaluating across four standard MRI sequences, our experiments demonstrate that the Sagittal T2-weighted sequence offers the most robust diagnostic value, achieving a superior Macro F1-score of 50.31%. We find that multi-sequence fusion strategies yield inferior performance compared to this single-sequence baseline, as the images across sequences in our dataset are significantly compromised by noise interference from surrounding anatomical regions. This work establishes a robust baseline and offers critical insights into sequence selection for spine analysis. 1 Introduction The human spine is a sophisticated anatomical axis, integrating osseous, ligamentous, and muscular components to ensure biomechanical stability and protect the central nervous system. Despite its structural robustness, the spinal column is susceptible to a heterogeneous array of pathologies. Currently, definitive diagnosis relies on a synthesis of clinical assessment and radiological imaging, often supplemented by invasive tissue sampling. However, biopsies are frequently hindered by sampling inaccuracies and the inherent risks of surgical complications [9, 7]. The temporal lag associated with histopathological turnaround can be detrimental to the patient’s oncological and functional prognosis [1]. Transitioning towards a purely Magnetic Resonance Imaging (MRI)-based diagnostic paradigm offers a significant opportunity to accelerate clinical workflows, mitigate procedural risks, and enhance global accessibility to specialized care. In recent years, the medical field has witnessed a transformative shift toward non-invasive, data-driven methodologies, with Deep Learning emerging as the predominant architecture for complex medical image analysis due to its superior hierarchical feature extraction capabilities [2, 4]. To bridge the gap between algorithmic potential and clinical utility, this paper provides several key contributions. (1) We introduce PhenSPINE: a comprehensive dataset specifically curated for spinal pathology diagnosis. (2) We develop and evaluate benchmark for four distinct MRI sequences, establishing a performance baseline across diverse imaging sequences. (3) We propose a Positional Encoding mechanism to explicitly model the anatomical location of intervertebral discs, significantly improving diagnostic performance by leveraging spatial context. (4) Our empirical analysis identifies the Sagittal T2-weighted (SAG-T2) sequence as the most informative sequence, yielding the highest diagnostic precision. (5) We demonstrate that integrating all four MRI sequences leads to performance degradation compared to the best-performing single-sequence model, highlighting noise in spinal image fusion in our dataset. The remainder of this paper is organized as follows. Section 2 reviews related work. Section 3 details the PhenSPINE dataset and the benchmark. Section 4 analyzes the experimental results. Finally, Section 5 concludes the paper. 2 Related Work 2.1 Spine Medical Imaging Datasets. Access to high-quality, annotated datasets is a prerequisite for developing robust computer-aided diagnosis systems. The UW-Spine dataset [3], comprised of CT scans from 125 patients, served as a foundational resource for vertebra localization, particularly in pathological cases involving scoliosis and metal implants. Following this, the Large Scale Vertebrae Segmentation Challenge datasets [13] established a gold standard for volumetric spine segmentation. These datasets provide diverse CT scans with voxel-level annotations and centroid coordinates, addressing challenges in diverse field-of-views and anatomical variations. In the domain of Magnetic Resonance Imaging, the Lumbar Spine MRI Dataset [15] has been widely utilized for intervertebral disc classification and Pfirrmann grading. Most recently, the RSNA 2024 Lumbar Spine Degenerative Classification competition [11] introduced a large-scale, multi-center MRI dataset labeled by expert neuroradiologists. Unlike previous datasets restricted to single pathologies, RSNA 2024 necessitates the simultaneous classification of multiple degenerative conditions-lumbar disc herniation, spinal canal stenosis, and neural foraminal narrowing-across five lumbar levels. In this work, we introduce PhenSPINE, a dataset designed to evaluate multi-sequence models for spinal diagnostics. 2.2 Deep Learning Approaches for Spine Pathology Diagnosis. Early efforts in spine analysis, particularly on datasets like UW-Spine, relied heavily on traditional machine learning techniques. Glocker et al. [3] employed regression forests for vertebrae localization, effectively handling pathological deformations without deep neural networks. However, the paradigm shifted with the advent of large-scale benchmarks. In segmentation tasks such as the VerSe challenge, a new performance standard was defined in [10] through the application of U-Net [12] and its variants within a coarse-to-fine framework. This approach first regresses a global heatmap to localize the spine and subsequently segments individual vertebrae with high precision. In the domain of pathology diagnosis on MRI, deep learning has replaced manual feature engineering. Natalia et al. [8] utilized these architectures to automate the classification of intervertebral discs and Pfirrmann grading. We are the first to apply positional encoding combined with a multi-sequence fusion approach for spine pathology diagnosis. In this work, we propose a diagnostic benchmark that integrates Positional Encoding to explicitly model the anatomical context of intervertebral discs. The specific details of these contributions, including the dataset characteristics and the proposed methodology, will be presented in the section below. 3 Materials and method 3.1 PhenSPINE Dataset 3.1.1 Data Collection. Figure 1: Overview of the PhenSPINE data collection process. A total of 16,813 raw DICOM images representing a cohort of 250 unique patients were curated at PHENIKAAMEC hospital and annotated by expert physicians, as illustrated in Fig. 1. To ensure temporal diversity and mitigate potential bias associated with longitudinal variations in imaging equipment or protocols, the benchmark was partitioned into three chronological subsets collected throughout 2025, as summarized in Table 1. Regarding spine pathology annotations, each IVD instance is specified by four independent binary labels, its anatomical location within the lumbar spine (IVD labels 1 to 5), and the integrated multi-sequence information comprising Sagittal T2-weighted (SAG-T2), Sagittal T1-weighted (SAG-T1), Axial T2-weighted (AX-T2), and Sagittal STIR (SAG-STIR). Table 1: Quantitative summary of the PhenikaaMed MRI subsets Subset Patients DICOM Files Collection Period T27.7.25 5151 3,4583,458 July 2025 T8 7070 4,5764,576 August 2025 T9 129129 8,7798,779 September 2025 Total 250 ,16,813 – 3.1.2 Data Distribution. The distribution of spinal pathologies within the PhenSPINE dataset exhibits a significant class imbalance, reflecting real-world clinical prevalence where pathological cases are less frequent than normal findings. As presented in Table 2, the dataset comprises a total of 1,185 annotated IVD records. Normal cases constitute the majority class, accounting for 66.41% of the population. Among the specific pathologies, Disc Bulging is the most prevalent condition, occurring in 26.75% of cases. In contrast, Spondylolisthesis and Disc Narrowing are relatively rare, appearing in only 3.29% and 4.73% of records, respectively. Disc Herniation is observed in 6.92% of the dataset. Fig. 2 illustrates the distinct morphological characteristics of these four pathology classes, as visualized in the SAG-T2 sequence. Table 2: Distribution of spinal pathologies in the PhenSPINE dataset Pathology Pos. Neg. Total Prev. (%) Herniation 82 1103 1185 6.92 Bulging 317 868 1185 26.75 Spondylolisthesis 39 1146 1185 3.29 Narrowing 56 1129 1185 4.73 Normal 787 398 1185 66.41 (a) Herniation (b) Bulging (c) Spondylolisthesis (d) Narrowing Figure 2: Representative MRI examples (SAG-T2) for the four analyzed categories: Disc Herniation, Disc Bulging, Spondylolisthesis, and Disc Narrowing. 3.1.3 Data Preprocessing. Missing sequences are handled via zero-padding to preserve consistent input dimensionality. A central slice is selected from each sequence using: Y=⌊n2⌋+1,Y= n2 +1, (1) where n is the total slice count and Y the selected index, ensuring the input corresponds to the anatomical midpoint where pathological features are most pronounced. Raw DICOM images are converted to PNG and resized to 224×224224× 224 pixels. During training, augmentation via torchvision includes geometric transforms (RandomHorizontalFlip, RandomRotation, RandomAffine with translation up to 15%15\% and scale [0.85,1.15][0.85,1.15]) and photometric transforms (ColorJitter, GaussianBlur σ∈[0.1,0.5]σ∈[0.1,0.5] with p=0.3p=0.3). Sequence-specific normalization uses adapted ImageNet statistics (μ=0.449μ=0.449, σ=0.226σ=0.226) for grayscale inputs. Each IVD is ultimately represented as a multi-sequence tensor ∈ℝ4×224×224X ^4× 224× 224. 3.2 PhenSPINE Benchmark Following the data pre-processing stage to ensure the integrity of the input features, we proposed a pipeline for benchmarking the multi-sequence fusion approach. This framework is specifically designed to evaluate the performance of integrating diverse MRI sequences for detecting complex spinal pathologies. Figure 3: Architecture of the PhenSPINE Benchmark. The pipeline comprises five key stages: (1) Data Pre-processing, (2) Feature Extraction, (3) Positional Encoding for Intervertebral Disc Metadata, (4) Anatomy-Aware Feature Fusion, and (5) Multi-Pathology Classification Head. 3.2.1 Feature Extraction. Given a collection of multi-sequence MRI scans Xii=14\X_i\_i=1^4, where each Xi∈ℝH×W×CX_i ^H× W× C corresponds to a specific MRI sequence, high-level spatial feature maps FiF_i are extracted according to: Fi=ℱ(Xi;θ),i∈1,2,3,4,F_i=F(X_i;θ), i∈\1,2,3,4\, (2) where ℱF denotes the feature extraction backbone parameterized by θ. A global average pooling operation is subsequently applied to each feature map FiF_i to aggregate local spatial information into a global descriptor: gi=GlobalAvgPool(Fi),i∈1,2,3,4,g_i=GlobalAvgPool(F_i), i∈\1,2,3,4\, (3) where gi∈ℝCg_i ^C represents the spatially-aggregated feature vector for the i-th MRI sequence. These global descriptors are then directly concatenated along the feature dimension to leverage complementary semantic information across all MRI sequences, forming a unified multi-sequence representation m: =[g1;g2;g3;g4],∈ℝ4C,m=[g_1;g_2;g_3;g_4], ^4C, (4) where C specifies the dimensionality of each individual sequence feature vector, and the semicolon notation denotes concatenation. 3.2.2 Positional Encoding for Intervertebral Disc Metadata. Inspired by the Transformer [16], we adapted the Positional Encoding mechanism to suit this domain. The categorical metadata specifying the IVD level for each input sequence is mapped to a discrete numerical representation through a label encoding function E. Let ℒ=l1,l2,…,l5L=\l_1,l_2,…,l_5\ denote the set of lumbar IVD levels, where each level is assigned a unique integer index according to: yivd=E(li),yivd∈0,1,2,3,4.y_ivd=E(l_i), y_ivd∈\0,1,2,3,4\. (5) The encoded index yivdy_ivd serves as an anatomical identifier that facilitates retrieval of the corresponding positional representation from a learned embedding space. To effectively encode the distinctive anatomical characteristics associated with each IVD level, we employ a trainable embedding matrix pos∈ℝN×dsW_pos ^N× d_s, where N=5N=5 represents the total number of lumbar levels and dsd_s specifies the dimensionality of the positional embedding space. This embedding mechanism projects the discrete level index yivdy_ivd into a continuous, high-dimensional representation according to: pos=Lookup(yivd,pos),pos∈ℝds,e_pos=Lookup(y_ivd,W_pos), _pos ^d_s, (6) where pose_pos denotes the resulting positional embedding vector. This learned representation encodes the spatial characteristics specific to each IVD level, thereby enabling the model to modulate its diagnostic inference according to the anatomical context of the input data. 3.2.3 Anatomy-Aware Feature Fusion. To establish a comprehensive representation integrating both multi-sequence features and anatomical context, the positional embedding pose_pos is combined with the global multi-sequence feature vector m. The anatomy-aware feature vector z is formulated through concatenation: =[;pos],∈ℝ4C+ds,z=[m;e_pos], ^4C+d_s, (7) where the semicolon denotes concatenation along the feature dimension. This joint representation z enhances model interpretability and diagnostic robustness across distinct lumbar levels. 3.2.4 Multi-Pathology Classification Head. To detect pathology, the fused anatomy-aware feature vector z is processed by a multi-head classification architecture comprising four independent feedforward networks. Each classification head j∈1,2,3,4j∈\1,2,3,4\ maps the shared representation z to a pathology-specific probability pjp_j through a two-layer nonlinear transformation followed by sigmoid activation: pj=σ(2,j⊤⋅ReLU(1,j⊤+1,j)+b2,j),p_j=σ (W_2,j ·ReLU(W_1,j z+b_1,j)+b_2,j ), (8) where 1,jW_1,j, 2,jW_2,j, 1,jb_1,j, and b2,jb_2,j denote the learnable parameters of the j-th classification head. The individual probabilities are subsequently assembled into a unified prediction vector =[p1,p2,p3,p4]⊤p=[p_1,p_2,p_3,p_4] . 3.2.5 Loss Function. We employed the Binary Cross-Entropy loss function for the multi-label classification task, formulated to aggregate the prediction errors across all pathology heads. The objective function is defined as: ℒ=−1N∑i=1N∑j=14[yj(i)log(pj(i))+(1−yj(i))log(1−pj(i))],L=- 1N _i=1^N _j=1^4 [y_j^(i) (p_j^(i))+(1-y_j^(i)) (1-p_j^(i)) ], (9) where N is the batch size, j indexes the classification heads, yj(i)y_j^(i) denotes the ground truth label, and pj(i)p_j^(i) represents the prediction for the i-th sample. 4 Experiments 4.1 Experimental Setup. We utilize the PhenSPINE dataset to analyze spinal pathology detection performance. To ensure robust evaluation, the data is chronologically partitioned into training (70%70\%), validation (15%15\%), and testing (15%15\%) subsets. All experiments were implemented using the PyTorch framework on an NVIDIA RTX 4090 GPU with 24 GB of memory. We employ a comprehensive set of metrics including Accuracy, Recall, Precision, Macro F1-score, and AUC to assess model performance. In our experimental tables, the best results are marked in bold, while the second-best results are underlined. 4.2 Comparative Evaluation of Feature Extraction Architectures. Table 3: Classification performance comparison across backbone architectures Backbone Sequence Acc Recall Pre Macro F1 AUC DenseNet [6] AX-T2 42.22 71.89 31.05 41.01 81.85 SAG-STIR 59.44 37.15 30.81 33.14 67.28 SAG-T1 57.22 50.69 38.20 40.98 82.43 SAG-T2 66.67 54.93 48.60 49.76 88.41 ResNet [5] AX-T2 57.78 56.19 37.50 42.77 82.66 SAG-STIR 56.67 43.89 47.49 32.35 70.08 SAG-T1 52.22 54.43 31.28 37.08 79.26 SAG-T2 60.00 62.16 40.22 42.37 85.75 EffB2 [14] AX-T2 56.11 67.64 38.05 46.22 85.67 SAG-STIR 55.00 46.28 24.97 32.04 73.07 SAG-T1 63.89 47.14 44.70 40.22 85.67 SAG-T2 54.44 64.65 43.64 47.74 87.69 EffB1 [14] AX-T2 60.56 70.11 36.33 47.16 86.50 SAG-STIR 56.67 40.27 25.14 30.11 74.45 SAG-T1 55.56 71.92 33.49 43.99 85.87 SAG-T2 57.78 59.87 60.01 50.31 87.46 EffB0 [14] AX-T2 58.33 55.85 39.71 43.39 84.15 SAG-STIR 60.00 48.60 28.01 34.68 71.02 SAG-T1 63.89 48.32 36.27 40.43 86.23 SAG-T2 66.11 44.46 48.04 43.02 88.47 Table 3 presents a comparative evaluation of five backbone architectures across four distinct MRI sequences. The empirical results indicate that the EfficientNet-B1 backbone, when paired with the SAG-T2 sequence, achieves the highest Macro F1-score of 50.31% and a Precision of 60.01%, establishing an optimal balance between discriminative power and computational efficiency. While DenseNet with SAG-T2 achieves the highest overall Accuracy (66.67%) and EfficientNet-B0 with SAG-T2 attains the peak AUC (88.47%), EfficientNet-B1 consistently demonstrates superior F1-scores . Conversely, larger architectures like EfficientNet-B2 and ResNet fail to surpass the performance of the more compact B1 variant, suggesting that increased model complexity does not necessarily yield enhanced diagnostic accuracy in this domain. From this evaluation, we derive a primary observation: EfficientNet-B1 demonstrates a superior ability to mitigate the overfitting often observed in deeper networks while retaining sufficient capacity to capture subtle pathological features. Consequently, we select the EfficientNet-B1 architecture as the primary backbone for subsequent multi-sequence fusion experiments. 4.3 Evaluation of Single-Sequence Models without IVD Metadata. Table 4: Diagnostic performance across different MRI sequences without IVD metadata Sequence Acc Recall Pre F1 AUC SAG-T1 2.78 60.81 23.06 26.96 66.13 SAG-T2 36.11 49.67 28.12 27.92 73.61 AX-T2 42.22 33.44 27.27 27.23 64.91 SAG-STIR 6.11 43.30 11.91 15.84 56.13 Table 4 details the diagnostic performance of backbone models when trained on individual MRI sequences without the inclusion of IVD metadata. The empirical evidence reveals that while T2-weighted modalities generally exhibit superior performance metrics-with SAG-T2 attaining the highest F1-score (27.92%) and AX-T2 reporting the highest Accuracy (42.22%)-all sequences struggle notably compared to models integrated with IVD metadata. These results underscore a fundamental characteristic of this diagnostic task: the pronounced performance degradation across all sequences highlights that IVD positional metadata is not merely an auxiliary feature but a prerequisite for enabling the model to localize and differentiate vertebral structures accurately. 4.4 Impact of Positional Encoding for Intervertebral Disc Metadata. Table 5: Comparison of IVD encoding methods Label Encoding Positional Encoding Sequence Acc Recall Pre F1 AUC Acc Recall Pre F1 AUC SAG-T1 63.33 58.87 37.72 44.61 86.87 55.56 71.92 33.49 43.99 85.87 SAG-T2 68.33 62.85 39.58 47.60 87.56 57.78 59.87 60.01 50.31 87.46 AX-T2 58.89 60.83 40.90 44.99 87.40 60.56 70.11 36.33 47.16 86.50 SAG-STIR 58.33 37.44 29.53 32.91 75.52 56.67 40.27 25.14 30.11 74.45 The comparative results presented in Table 5 quantify the significant contribution of Positional Encoding for IVD metadata. Specifically, for the optimal SAG-T2 sequence, while standard Label Encoding yields a higher raw Accuracy of 68.33%68.33\%, it is compromised by a substantially lower Precision (39.58%39.58\%) and Macro F1-score (47.60%47.60\%). In contrast, the integration of Positional Encoding elevates Precision to 60.01%60.01\% and the F1-score to 50.31%50.31\%. These empirical observations underscore the efficacy of Positional Encoding as a spatial inductive bias essential for spinal pathology localization. Positional Encoding maps IVD levels into a continuous vector space, thereby capturing relative spatial distances and sequential dependencies between levels. This structure-aware representation effectively guides the backbone network to differentiate between similar visual patterns at distinct anatomical locations, transitioning the diagnostic process from image classification toward a more context-aware clinical assessment. 4.5 Multi-Sequence Fusion Strategies. Table 6: Classification performance on different sequence MRI combinations Methods Acc Recall Pre F1 Auc SAG-T1 55.56 71.92 33.49 43.99 85.87 SAG-T2 57.78 59.87 60.01 50.31 87.46 SAG-STIR 56.67 40.27 25.14 30.11 74.45 AX-T2 60.56 70.11 36.33 47.16 86.50 SAG-T1+AX-T2 56.11 48.16 40.73 42.05 78.24 SAG-T1+SAG-T2 64.44 67.48 39.42 48.74 88.84 SAG-T1+SAG-STIR 63.33 73.98 38.35 49.44 89.30 SAG-T2+AX-T2 55.00 51.29 52.06 44.29 87.47 SAG-T2+SAG-STIR 54.44 58.70 38.88 44.77 86.65 AX-T2+SAG-STIR 58.89 54.27 30.65 37.83 84.69 SAG-T1+SAG-T2+AX-T2 65.00 49.51 45.19 45.27 88.94 SAG-T1+SAG-T2+SAG-STIR 63.89 52.16 42.32 45.96 87.03 SAG-T1+AX-T2+SAG-STIR 58.89 56.77 41.09 45.74 85.60 SAG-T2+AX-T2+SAG-STIR 58.89 60.35 39.69 45.07 86.73 All Sequences 67.22 61.63 40.37 48.33 88.86 Table 6 presents the classification performance across various MRI sequence combinations. The SAG-T2 sequence demonstrates superior performance across critical discriminative metrics, enabling an efficient diagnostic model. Specifically, SAG-T2 achieves a Macro F1-score of 50.31%50.31\% and a Precision of 60.01%60.01\%, substantially outperforming the multi-sequence fusion models and other single sequences. Notably, while the All Sequences fusion strategy achieves the highest global Accuracy (67.22%67.22\%) and the combination of SAG-T1 and SAG-STIR yields the peak AUC (89.30%89.30\%), these gains do not translate to better pathology-specific detection as evidenced by lower F1-scores (48.33%48.33\% and 49.44%49.44\% respectively). The comparative analysis highlights the specialized efficacy of single-sequence feature learning for spinal diagnostics. Although multi-sequence integration enhances broad performance metrics such as global Accuracy and AUC, the SAG-T2 sequence possesses high information density, enabling the network to distill essential diagnostic markers without noise introduced by auxiliary sequences. Therefore, the combination of SAG-T2 and EfficientNet-B1 is established as the optimal configuration, offering a clinically interpretable benchmark for characterizing spinal pathologies. 5 Conclusion In this study, we introduced PhenSPINE, a comprehensive MRI dataset designed to advance spinal pathology diagnosis. By curating a diverse dataset and establishing a rigorous evaluation protocol, we analyzed the efficacy of different MRI sequences and deep learning architectures. Our benchmark, which integrates a Positional Encoding mechanism, demonstrated the critical importance of anatomical context, significantly outperforming standard label encoding approaches. Extensive experiments revealed that the SAG-T2-weighted sequence provides the most discriminative features for pathology detection, whereas multi-sequence fusion strategies failed to yield performance gains due to feature redundancy and noise in our dataset. To address these limitations, future work will focus on incorporating explicit spine segmentation into the pipeline to precisely isolate vertebral structures, thereby eliminating background noise and facilitating more effective multi-sequence integration. Finally, to support the biomedical research community and foster further innovation, we are committed to making the PhenSPINE dataset publicly available on an appropriate open-access platform. Acknowledgement This research is supported by the Ben Dam Me Award Fund, the Vietnam Young Talent Support Fund, and the Number One Brand, Tan Hiep Phat Group. References [1] S. Alshieban and K. Al-Surimi (2015) Reducing turnaround time of surgical pathology reports in pathology and laboratory medicine departments. BMJ Qual. Improv. Rep. 4 (1). Cited by: §1. [2] A. Esteva, A. Robicquet, B. Ramsundar, V. Kuleshov, M. DePristo, K. Chou, C. Cui, G. Corrado, S. Thrun, and J. Dean (2019) A guide to deep learning in healthcare. Nat. Med. 25 (1), p. 24–29. Cited by: §1. [3] B. Glocker, D. Zikic, E. Konukoglu, D. R. Haynor, and A. Criminisi (2013) Vertebrae localization in pathological spine CT via dense classification from sparse annotations. In MICCAI, p. 262–270. Cited by: §2.1, §2.2. [4] R. Grossman, O. Haim, S. Abramov, B. Shofty, and M. Artzi (2021) Differentiating small-cell lung cancer from non-small-cell lung cancer brain metastases based on MRI using EfficientNet and transfer learning approach. TCRT 20, p. 15330338211004919. Cited by: §1. [5] K. He, X. Zhang, S. Ren, and J. Sun (2016) Deep residual learning for image recognition. In CVPR, p. 770–778. Cited by: Table 3. [6] G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger (2017) Densely connected convolutional networks. In CVPR, p. 4700–4708. Cited by: Table 3. [7] R. Nasser, S. Yadla, M. G. Maltenfort, J. S. Harrop, D. G. Anderson, A. R. Vaccaro, A. D. Sharan, and J. K. Ratliff (2010) Complications in spine surgery: a review. J. Neurosurg. Spine 13 (2), p. 144–157. Cited by: §1. [8] F. Natalia, S. Sudirman, D. Ruslim, and A. Al-Kafri (2024) Lumbar spine MRI annotation with intervertebral disc height and pfirrmann grade predictions. PLOS ONE 19 (5), p. e0302067. Cited by: §2.2. [9] A. Olscamp, J. Rollins, S. S. Tao, and N. A. Ebraheim (1997) Complications of CT-guided biopsy of the spine and sacrum. Orthopedics 20 (12), p. 1149–1152. Cited by: §1. [10] C. Payer, D. Stern, H. Bischof, and M. Urschler (2020) Coarse to fine vertebrae localization and segmentation with spatialconfiguration-net and u-net.. In VISIGRAPP (5: VISAPP), p. 124–133. Cited by: §2.2. [11] T. J. Richards, A. E. Flanders, E. Colak, L. M. Prevedello, R. L. Ball, F. Kitamura, J. Mongan, M. Vazirabad, H. Lin, A. Kendell, et al. (2026) The RSNA lumbar degenerative imaging spine classification (LumbarDISC) dataset. Radiology: Artificial Intelligence, p. e250480. Cited by: §2.1. [12] O. Ronneberger, P. Fischer, and T. Brox (2015) U-Net: convolutional networks for biomedical image segmentation. In MICCAI, p. 234–241. Cited by: §2.2. [13] A. Sekuboyina, M. E. Husseini, A. Bayat, M. Löffler, H. Liebl, H. Li, G. Tetteh, J. Kukačka, C. Payer, D. Štern, et al. (2021) VerSe: a vertebrae labelling and segmentation benchmark for multi-detector CT images. Med. Image Anal. 73, p. 102166. Cited by: §2.1. [14] M. Tan and Q. Le (2019) EfficientNet: rethinking model scaling for convolutional neural networks. In ICML, p. 6105–6114. Cited by: Table 3, Table 3, Table 3. [15] J. W. van der Graaf, M. L. van Hooff, C. F. Buckens, M. Rutten, J. L. van Susante, R. J. Kroeze, M. de Kleuver, B. van Ginneken, and N. Lessmann (2024) Lumbar spine segmentation in MR images: a dataset and a public benchmark. Sci. Data 11 (1), p. 264. Cited by: §2.1. [16] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin (2017) Attention is all you need. NeurIPS 30. Cited by: §3.2.2.