Paper deep dive
Longitudinal Multi-View Breast Cancer Risk Prediction
Solveig Thrun, Zijun Sun, Suaiba A. Salahuddin, Kristoffer Wickstrøm, Elisabeth Wetzer, Stine Hansen, Robert Jenssen, Michael Kampffmeyer
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/18/2026, 2:28:32 PM
Summary
The paper introduces LMV-Net, a deep learning model for longitudinal multi-view breast cancer risk prediction. It addresses limitations in existing methods by explicitly aligning CC and MLO mammographic views across time points using deformation fields and a dual-stream attention mechanism. Evaluated on EMBED and CSAW-CC datasets, LMV-Net outperforms state-of-the-art baselines in C-index and AUC, demonstrating improved risk stratification, particularly for dense breast tissue and invasive cancer subtypes.
Entities (10)
Relation Signals (13)
LMV-Net → evaluatedon → EMBED
confidence 95% · We evaluate our approach on the public EMBED and CSAW-CC datasets
LMV-Net → evaluatedon → CSAW-CC
confidence 95% · We evaluate our approach on the public EMBED and CSAW-CC datasets
LMV-Net → uses → Dual Stream Attention
confidence 95% · We propose a Dual Stream Attention block designed to enhance the modeling of longitudinal multi-view features.
LMV-Net → processes → MLO
confidence 92% · jointly analyzes anatomically complementary CC and MLO views
LMV-Net → processes → CC
confidence 92% · jointly analyzes anatomically complementary CC and MLO views
LMV-Net → measuredby → AUC
confidence 90% · we evaluate performance using the C-index and AUC
LMV-Net → measuredby → C-index
confidence 90% · we evaluate performance using the C-index and AUC
LMV-Net → outperforms → VMRA-MaR
confidence 90% · LMV-Net consistently outperforms all baselines in C-index and AUC
LMV-Net → outperforms → ImgFeatAlign
confidence 90% · LMV-Net consistently outperforms all baselines in C-index and AUC
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Accurate breast cancer risk prediction from screening mammography is critical for enabling personalized screening intervals and early detection. Recent deep learning methods have shown the value of longitudinal data and explicit temporal alignment. However, existing approaches either perform explicit alignment using a single mammographic view or model multiple views without explicit longitudinal alignment, limiting their ability to exploit the complementary spatial-temporal information used in clinical practice. To address this gap, we propose LMV-Net, a longitudinal multi-view breast cancer risk prediction model that jointly analyzes anatomically complementary CC and MLO views within an explicitly aligned longitudinal framework. We evaluate our approach on the public EMBED and CSAW-CC datasets, comparing it to state-of-the-art breast cancer risk prediction methods. Our model consistently outperforms existing approaches in overall risk prediction performance and across different breast density and cancer subgroups. Importantly, these improvements highlight the potential of longitudinal multi-view modeling to enhance risk stratification, paving the way for future work on personalized screening, earlier identification of high-risk patients, and more efficient screening resource allocation. The code is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2607.11343v1
- Canonical: https://arxiv.org/abs/2607.11343v1
Trouble viewing inline? Open PDF directly →
Full Text
32,263 characters extracted from source content.
Expand or collapse full text
11institutetext: Department of Physics and Technology, UiT The Arctic University of Norway, Tromsø, Norway 11email: solveig.thrun@uit.no 22institutetext: SPKI The Norwegian Centre for Clinical Artificial Intelligence, University Hospital of North Norway, Tromsø, Norway 33institutetext: Norwegian Computing Center, Oslo, Norway 44institutetext: Pioneer Centre for AI, University of Copenhagen, Copenhagen, Denmark Longitudinal Multi-View Modeling for Breast Cancer Risk Prediction Solveig Thrun Zijun Sun Suaiba A. Salahuddin Kristoffer Wickstrøm Elisabeth Wetzer Stine Hansen Robert Jenssen Michael Kampffmeyer Abstract Accurate breast cancer risk prediction from screening mammography is critical for enabling personalized screening intervals and early detection. Recent deep learning methods have shown the value of longitudinal data and explicit temporal alignment. However, existing approaches either perform explicit alignment using a single mammographic view or model multiple views without explicit longitudinal alignment, limiting their ability to exploit the complementary spatial–temporal information used in clinical practice. To address this gap, we propose LMV-Net, a longitudinal multi-view breast cancer risk prediction model that jointly analyzes anatomically complementary C and MLO views within an explicitly aligned longitudinal framework. We evaluate our approach on the public EMBED and CSAW-C datasets, comparing it to state-of-the-art breast cancer risk prediction methods. Our model consistently outperforms existing approaches in overall risk prediction performance and across different breast density and cancer subgroups. Importantly, these improvements highlight the potential of longitudinal multi-view modeling to enhance risk stratification, paving the way for future work on personalized screening, earlier identification of high-risk patients, and more efficient screening resource allocation. The code is available at https://github.com/sot176/LMV-Net. 1 Introduction Breast cancer is the most common cancer among women worldwide, representing a major public health challenge [5]. Early detection through population-based screening has been critical in reducing mortality, with mammography, typically acquired in craniocaudal (C) and mediolateral oblique (MLO) views, remaining the gold standard [3, 30]. Large-scale screening programs have therefore played a key role in identifying cancers at an earlier, better treatable stage [30]. However, effective risk stratification for personalized screening and early detection in high-risk populations remains limited [21], especially in women with dense breast tissue, where density both increases cancer risk and reduces mammographic sensitivity [4, 27]. Recent advances in deep learning have improved breast cancer risk prediction from mammography, offering a data-driven path toward more individualized screening [13]. Early deep learning models analyzed current images in isolation (e.g., Mirai [29]), overlooking the valuable longitudinal information that radiologists exploit when comparing current and prior examinations [1, 26]. To leverage this longitudinal information, recent models incorporate prior screenings, enabling them to capture gradual tissue changes over time and improve risk assessment [8, 15, 16, 18, 22, 28]. These models consistently outperform single-time-point approaches, demonstrating the value of modeling temporal information. Building on advances in mammography research, most longitudinal mammography models rely on implicit alignment (e.g., LoMaR [16], VMRA-MaR [22]), comparing prior and current images without enforcing spatial correspondence. Although non-rigid breast deformation makes explicit alignment challenging, recent evidence indicates that such strategies are better suited to capture subtle temporal variations, outperforming implicit approaches for modeling longitudinal change [23]. Nevertheless, existing explicit alignment methods (ImgFeatAlign [23, 24], OA-BreaCR [28]) have been restricted to single-view formulations, without jointly modeling the complementary information available across C and MLO views. As a result, current approaches do not fully exploit the combined temporal and cross-view cues that are critical for accurate breast cancer risk prediction. To address this shortcoming, we introduce a breast-wise longitudinal multi-view risk prediction framework that explicitly aligns the feature maps of each C and MLO view across multiple time points, capturing subtle longitudinal changes and more closely mirroring radiologists’ workflow. Our approach uses image-based deformation fields to align current and prior feature maps for each view, mitigating spatial misalignment while preserving discriminative information. By introducing a dual-stream attention mechanism, the model enables multi-view, breast-wise comparison over time. This design effectively captures longitudinal tissue changes and provides a foundation for more reliable risk prediction. Ultimately, it supports earlier identification of high-risk patients and facilitates personalized screening strategies. Our contributions are: 1. LMV-Net, a breast-wise longitudinal multi-view model for breast cancer risk prediction, that facilitates multi-view, multi-time-point integration via explicit feature-space alignment and cross-view temporal fusion. 2. A Dual Stream Attention block that refines longitudinal features of each view (C and MLO) using parallel self- and cross-attention. 3. Extensive validation on two large public datasets, EMBED and CSAW-C, demonstrating robust performance. 2 Method Figure 1: Schematic overview of the proposed LMV-Net. We aim to predict an individual’s future breast cancer risk over a five-year horizon using mammograms from current and prior screening visits for a single breast, to support personalized, risk-based screening programs. Unlike prior approaches that rely on explicit alignment and use only a single view [23, 28], our proposed LMV-Net (Fig. 1) incorporates both C and MLO views, which provide complementary information. To achieve this, we introduce a dual-stream attention block for multi-view fusion that refines explicitly aligned longitudinal features through parallel self- and cross-attention. LMV-Net comprises six sequential stages. Step 1: Feature extraction. A shared image encoder extracts feature maps from current and prior mammograms in the MLO and C views, yielding four feature maps: current feature maps curMLO,curCC∈ℝC×H×Wf_cur ,f_cur ^C× H× W and prior feature maps priMLO,priCC∈ℝC×H×Wf_pri ,f_pri ^C× H× W, where C is the number of channels, and H and W denote the height and width of the feature maps, respectively. Step 2: Temporal feature alignment. Inspired by [23], we explicitly align longitudinal feature maps in feature space using image-based deformation fields. The deformation field is obtained using a pretrained and frozen MammoRegNet [24] by registering the prior mammogram to the current mammogram. A separate deformation field is estimated for each view (C and MLO). Using a spatial transformer function [24, 28], the prior feature maps (priMLOf_pri and priCCf_pri ) are aligned to the corresponding current feature maps (curMLOf_cur and curCCf_cur ), resulting in the aligned prior feature maps pri-alignMLOf_pri-align and pri-alignCCf_pri-align . Step 3: Temporal subtraction. To capture temporal changes, we follow [23, 28] and compute view-specific difference features via temporal subtraction: diffCC=curCC−pri-alignCC,diffMLO=curMLO−pri-alignMLOf_diff =f_cur -f_pri-align , _diff =f_cur -f_pri-align . These difference features encode longitudinal changes between aligned current and prior mammograms, and positional encoding [28] is applied to diffCCf_diff and diffMLOf_diff to account for variable time intervals between screenings. Step 4: Temporal feature concatenation. Current, aligned prior, and difference features are concatenated along the channel dimension to form longMLO∈ℝ3C×H×Wf_long ^3C× H× W and longCC∈ℝ3C×H×Wf_long ^3C× H× W, integrating current appearance, longitudinal context, and temporal changes. Step 5: Dual Stream Attention. We propose a Dual Stream Attention block designed to enhance the modeling of longitudinal multi-view features. The module jointly refines the longitudinal representations longMLOf_long and longCCf_long by explicitly capturing both intra-view and inter-view dependencies. For each view, the block applies multi-head self-attention to model long-range intra-view interactions and improve spatial coherence within the same mammographic view. In parallel, multi-head cross-attention leverages features from the complementary view as keys and values, enabling each view to selectively attend to anatomically and semantically relevant regions in the other view. This mechanism facilitates effective information exchange between C and MLO representations, enabling the model to leverage complementary perspectives of the same breast tissue. The outputs of self- and cross-attention are fused via element-wise summation, after which stochastic depth (DropPath [12]) is applied to improve generalization, and reduce overfitting. This attention output is added to the original residual features, followed by layer normalization [2]. Each view is subsequently processed through a feed-forward network (FFN) with a fourfold expansion of the channel dimension, increasing model expressiveness. A second residual connection and normalization are applied after the FFN. Through this dual-stream attention design, each view integrates both its own global spatial context and complementary cross-view information, resulting in enriched and more discriminative multi-view feature representations. Step 6: Final fusion and risk prediction. The refined longMLOf_long and longCCf_long features are pooled and concatenated into a fused multi-view representation f, which passes through a cumulative probability layer [18, 29] for time-dependent risk estimation. This layer comprises two fully connected layers: one predicting hazards at each future time step (i.e. 1 year, 2 years, …) and another computing a baseline hazard, with ReLU enforcing non-negativity. The cumulative probability at time t is then computed as P(t∣)=hbase()+∑j=1thj(),P(t )=h_base(f)+ _j=1^th_j(f), (1) where hbase()h_base(f) denotes the baseline hazard and hj()h_j(f) the hazard at time step j, capturing the sequential accumulation of risk. Training uses a time-dependent binary cross-entropy (BCE) loss, with labels set to 1 if cancer occurs within t years and 0 otherwise. For censored patients, a masking function δ(t)δ(t) [28, 29] is applied so that only valid time points contribute to the summed BCE loss,. Following [6, 28], we adopt multilevel joint learning to capture complementary dependencies between C and MLO views. The network jointly learns a fused multi-view representation for risk estimation and view-specific features that serve as auxiliary supervision during training, encouraging discriminative per-view learning and enhancing the quality of the fused representation used at inference. 3 Experimental Setup and Results Table 1: Breast-level C+MLO exam splits and time-to-cancer labels. EMBED Positive Negative Split 1Y 2Y 3Y 4Y 5Y All >>5Y Train 370 96 83 46 34 629 12262 Val 146 32 29 20 14 241 4907 Test 220 76 54 37 29 416 7341 CSAW- C Positive Negative Split 1Y 2Y 3Y 4Y 5Y All >>5Y Train 352 350 94 224 88 1108 15686 Val 140 142 48 92 28 450 6188 Test 208 214 76 124 54 676 9402 Datasets and Pre-processing We evaluate LMV-Net on two large, publicly available mammography datasets: the Emory Breast Imaging Dataset (EMBED) [14] and the Cohort of Screen-Aged Women Case Control dataset (CSAW-C) [9], ensuring reproducibility and an experimental setup mirroring prior work [16, 22, 23, 24, 28, 29]. EMBED comprises mammograms from a diverse population acquired on Hologic, GE, and Fujifilm systems, while CSAW-C contains cases collected at Karolinska University Hospital (Sweden) using Hologic systems. Following [28], we include only patients with at least five years of follow-up. All standard 2D mammograms are resized to 1664×20481664× 2048 pixels while preserving aspect ratio. Data are split at the patient level into training, validation, and test sets (5:2:3). CSAW-C provides LIBRA-based breast density estimates [17], grouped into three equally sized categories (low, medium, high), whereas EMBED uses BI-RADS density categories [7]. Time-to-cancer is defined as the interval between image acquisition and diagnosis. Table 1 summarizes dataset splits and time-to-cancer distributions at the breast-wise C–MLO exam level. For patients with more than two time points, we construct all consecutive time-point pairs of images for training and evaluation. For training MammoRegNet, we exclusively used patients who were excluded from the risk prediction cohort, preventing any potential data leakage. Evaluation Metrics Following [18, 28, 29], we evaluate performance using the C-index [25] and AUC [10] over 1–5 year risk horizons. 95% confidence intervals (CIs) are estimated via exam-level bootstrapping with 1,000 resamples. Statistical significant improvements over the second-best model is assessed using a paired non-parametric bootstrap test, with p<0.05p<0.05 considered significant. Implementation Details All models were implemented in PyTorch 2.0.1 [20] and on eight AMD Instinct MI210 GPUs (64 GB each) with a batch size of 4 per GPU (32 total). For fair comparison, we use the Mirai encoder [29], a ResNet-18 [11] with the final fully connected layer removed, following prior work [16, 22, 23, 28]. Self- and cross-attention use four heads, and training uses AdamW [19] with a learning rate of 5×10−55× 10^-5 and weight decay of 1×10−41× 10^-4. LMV-Net was trained for up to 30 epochs, with early stopping after 15 stagnant epochs based on the validation C-index and learning rate scheduling that halved the learning rate after 5 stagnant epochs. Data augmentation included random cropping, affine transformations, color jitter, and gamma correction. Baselines [22, 24, 28, 29] and MammoRegNet [23, 24] followed official implementations and hyperparameters, and LMV-Net was evaluated with both frozen and fine-tuned pretrained Mirai encoders. 3.1 Results Table 2: 1–5 Year breast cancer risk prediction: C-index and AUC values with ± 95% CI. : frozen encoder, : fine-tuned encoder. Δ denotes % improvement of best LMV-Net over best baseline. ∗ indicates statistically significant improvement of LMV-Net over the best baseline (p < 0.05). Training time in hours (h) and total number of model parameters (#, in millions). C-index (%) ↑ Follow-up year AUC (%) ↑ h # 1-Y 2-Y 3-Y 4-Y 5-Y EMBED Mirai [29] ( ) 59.8± 8.7 56.7± 20.7 52.6± 15.3 54.0± 12.6 49.3± 12.5 49.2± 11.6 6 28 Mirai [29] ( ) 69.1± 10.6 71.0± 15.6 70.9± 15.5 70.6± 11.7 65.7± 11.3 65.9± 11.6 7 28 VMRA-MaR [22] ( ) 72.4± 12.5 77.9± 4.5 75.3± 9.7 73.1± 12.7 68.9± 14.3 67.1± 11.8 45 2802 OA-BreaCR [28] ( ) 72.3± 2.0 72.9± 2.8 72.3± 4.9 72.4± 2.5 72.3± 2.4 73.2± 2.4 9 15 ImgFeatAlign [23] ( ) 74.7± 2.4 75.0± 2.9 75.5± 2.4 75.3± 2.1 75.9± 2.3 72.5± 3.2 12 52 LMV-Net ( ) 78.3∗± 3.0 80.8∗± 3.7 78.3∗± 3.5 77.5∗± 3.3 76.6± 3.2 75.1∗± 4.9 15 65 LMV-Net ( ) 81.4∗± 2.8 83.9∗± 3.4 82.1∗± 3.0 80.8∗± 2.9 80.0∗± 2.8 77.2∗± 4.7 16 65 Δ (%) ↑ + 6.7 + 6.0 + 6.6 + 5.5 + 4.1 + 4.0 CSAW-C Mirai [29]( ) 71.4± 2.7 56.8± 10.6 54.7± 5.8 56.4± 5.1 56.1± 4.2 55.4± 4.0 7 28 Mirai [29] ( ) 72.3± 2.9 69.9± 3.3 69.8± 3.9 69.2± 3.8 67.0± 3.4 66.9± 3.4 8 28 VMRA-MaR [22] ( ) 72.7± 2.6 70.9± 5.5 70.3± 5.6 69.5± 4.4 70.9± 4.3 64.5± 6.0 46 2802 OA-BreaCR [28] ( ) 61.6± 4.2 63.5± 4.4 61.7± 2.1 65.0± 2.0 67.0± 1.9 67.7± 2.2 10 15 ImgFeatAlign [23] ( ) 70.4± 1.9 72.0± 2.4 71.5± 1.9 72.6± 1.9 72.0± 2.1 75.2± 2.3 13 52 LMV-Net ( ) 73.4∗± 2.5 77.5∗± 3.3 73.5∗± 2.6 74.0± 2.5 73.1± 2.6 75.9± 2.8 16 65 LMV-Net ( ) 74.4∗± 2.7 78.7∗± 3.6 74.1∗± 2.9 74.3∗± 2.7 73.5± 2.8 75.0± 3.1 17 65 Δ (%) ↑ + 1.7 + 6.7 + 2.6 + 1.7 + 1.5 + 0.7 Comparison with State-of-the-Arts (SOTA) Table 2 compares the proposed LMV-Net with SOTA methods, including Mirai [29] (Sci. Transl. Med.’2021), OA-BreaCR [28] (MICCAI’24), VMRA-MaR [22] (MICCAI’25), and ImgFeatAlign [23] (MICCAI’25), on both datasets. LMV-Net consistently outperforms all baselines in C-index and AUC across all follow-up years on both datasets. Fine-tuning the encoder yields further improvements, achieving statistically significant gains (p<0.05p<0.05) over the second-best model. These results highlight the benefit of jointly modeling multi-view and longitudinal information for breast cancer risk prediction. Notably, the larger gains observed on EMBED are due to its greater heterogeneity across populations and imaging systems, whereas the single-scanner CSAW-C cohort is more homogeneous, limiting the potential benefits of improved model generalization. C-index by density category Higher breast density is associated with increased cancer risk [4] and reduced mammographic sensitivity [27]. Fig 2 (left) shows C-index values with 95% CI across density categories for EMBED and CSAW-C. Across both datasets, LMV-Net ( ) achieves the highest C-index across nearly all density groups. Note, some model estimates are not shown for high-density EMBED cases (category D) due to insufficient comparable pairs for reliable estimation. Gains are most pronounced in medium- and high-density groups, highlighting LMV-Net’s robustness in challenging screening scenarios. Figure 2: C-index by density category (left; EMBED and CSAW-C) and by cancer type (middle; EMBED) with 95% CI, and AUC over follow-up years for invasive cancer (right; EMBED). : frozen encoder, : fine-tuned encoder. Performance across cancer types We also perform the first systematic comparison of recent state-of-the-art breast cancer risk prediction methods across different cancer types on the more diverse EMBED dataset. As shown in Fig. 2 (middle), LMV-Net outperforms SOTA methods in C-index for both types, with the fine-tuned encoder achieving the best results and competing methods exhibiting greater uncertainty. Fig. 2 (right) shows AUC over follow-up years for invasive cancer, the more aggressive and clinically relevant type. LMV-Net achieves higher and more stable AUCs, whereas other methods show lower performance and greater uncertainty. These findings demonstrate that integrating multiview and longitudinal information in alignment with clinical workflow yields more robust and reliable breast cancer risk prediction, particularly for invasive disease. Attention map analysis Fig. 3 (left) visualizes the self- and cross-attention maps for a representative patient. LMV-Net assigns elevated attention to the tumor region (red bounding box) in both C and MLO views. Self-attention primarily focuses on the lesion and its surrounding area within each view, capturing intra-view dependencies. Cross-attention, integrates information from the complementary view, where we see the model attending primarily to the tumor tissue and/or suspicious regions in the other view. We further compute the ratio of mean attention density within the tumor bounding box to that within the entire breast region, aggregated across the 12 test-set patients with tumor annotations. (Fig. 3, right). On average, attention density within the tumor region is 8–10× higher than in the surrounding breast tissue. This supports that the proposed Dual Stream Attention block concentrates attention on anatomically corresponding tumor regions. Figure 3: Left: Attention maps showing LMV-Net focus on the tumor (red box). Right: Attention density in tumor bounding box relative to the whole breast. Ablation study We assess key architectural components through ablation studies: (i) Repeating current features: both temporal alignment (Step 2) and subtraction (Step 3) are removed, and Step 4 concatenates the current features three times; (i) Concatenating with zeros: temporal alignment and subtraction are removed as above, but Step 4 concatenates the current features with two zero tensors; (i) Single-view modeling: one branch of the model (C or MLO) is removed, and the cross-attention block in the remaining branch is replaced with an additional self-attention block. Only a single view is processed, effectively disabling multi-view input while preserving the same number of attention layers. (iv) Duplicated-view modeling: the same view is fed into both branches; (v) Dual-stream attention removed: the dual-stream attention block is removed; (vi) Multi-view learning: using only Ymulti-viewY_multi-view; and (vii) Encoder fine-tuning: freezing the pretrained Mirai backbone. Table 3 shows that removing longitudinal information, multi-view modeling, attention, or encoder fine-tuning consistently reduces C-index and AUC. The full model achieves the best performance on both EMBED and CSAW-C, confirming the value of each component. Table 3: Ablation study for 1–5 year risk: mean C-index and AUC on EMBED and CSAW-C. MV: Multi-view, DV: Different views in both branches, FT: Encoder fine-tuning, DSA: Dual Stream Attention. Variant Prior MV DV FT DSA YMLO_MLO YCC_C C-index (%) ↑ Follow-up year AUC (%) ↑ 1-Y 2-Y 3-Y 4-Y 5-Y EMBED (i) No prior (current repeated) ✗ ✓ ✓ ✓ ✓ ✓ 77.8 80.3 78.1 76.9 76.3 73.6 (i) No prior (zero-padded) ✗ ✓ ✓ ✓ ✓ ✓ 77.5 80.2 78.3 76.4 76.2 73.8 (i) Only one view ✓ ✗ ✓ ✓ ✓ ✓ 70.8 72.2 70.8 68.9 65.9 61.9 (iv) Same view both branches ✓ ✓ ✗ ✓ ✓ ✓ 77.4 78.8 78.1 77.7 76.2 74.7 (v) Without DSA ✓ ✓ ✓ ✓ ✗ ✓ 72.8 73.9 73.7 72.8 70.3 65.8 (vi) Use only Ymulti-view_multi-view ✓ ✓ ✓ ✓ ✓ ✗ 80.1 83.0 80.4 79.4 78.8 76.5 (vii) No finetuning ✓ ✓ ✓ ✗ ✓ ✓ 78.3 80.8 78.3 77.5 76.6 75.1 Ours ✓ ✓ ✓ ✓ ✓ ✓ 81.4 83.9 82.1 80.8 80.0 77.2 CSAW-C (i) No prior (current repeated) ✗ ✓ ✓ ✓ ✓ ✓ 73.0 75.3 73.8 73.2 72.3 73.2 (i) No prior (zero-padded) ✗ ✓ ✓ ✓ ✓ ✓ 72.2 75.2 73.0 72.6 71.9 72.4 (i) Only one view ✓ ✗ ✓ ✓ ✓ ✓ 73.3 73.8 73.1 73.5 73.1 73.9 (iv) Same view both branches ✓ ✓ ✗ ✓ ✓ ✓ 72.2 75.2 72.6 72.5 71.1 73.0 (v) Without DSA ✓ ✓ ✓ ✓ ✗ ✓ 69.5 70.5 70.0 70.6 69.3 71.4 (vi) Use only Ymulti-view_multi-view ✓ ✓ ✓ ✓ ✓ ✗ 73.3 77.6 73.3 73.4 72.6 74.4 (vii) No finetuning ✓ ✓ ✓ ✗ ✓ ✓ 73.4 77.5 73.5 74.0 73.1 75.9 Ours ✓ ✓ ✓ ✓ ✓ ✓ 74.4 78.7 74.1 74.3 73.5 75.0 4 Conclusion and Outlook In this study, we introduced LMV-Net, which mimics the radiologist’s workflow by jointly modeling longitudinal and multi-view information through explicit longitudinal alignment. Experiments on two large publicly available datasets demonstrate statistically significant improvements over SOTA methods, particularly in challenging scenarios such as patients with higher breast density. Clinically, this approach could improve early detection and personalized screening through more accurate risk assessment. While our results demonstrate the value of joint longitudinal multi-view modeling, further investigation is warranted. Incorporating multiple prior examinations may improve performance by providing richer temporal context. In addition, our framework relies on explicit feature-space alignment to enable accurate longitudinal modeling such that advances in robust explicit alignment methods could further enhance its performance. Future work will also focus on more comprehensive evaluation across additional clinically relevant metrics, including precision–recall AUC (PR-AUC), particularly given the class imbalance characteristic of breast cancer risk prediction. References [1] J. Akwo, P. Y. Trieu, M. Barron, T. Reynolds, and S. Lewis (2024) Access to prior screening mammograms affects the specificity but not sensitivity of radiologists’ performance. Clinical Radiology 79 (12). Cited by: §1. [2] J. L. Ba, J. R. Kiros, and G. E. Hinton (2016) Layer normalization. arXiv preprint arXiv:1607.06450. Cited by: §2. [3] A. Bhushan, A. Gonsalves, and J. U. Menon (2021) Current state of breast cancer diagnosis, treatment, and theranostics. Pharmaceutics 13 (5). Cited by: §1. [4] F.T.H. Bodewes, A.A. van Asselt, M.D. Dorrius, M.J.W. Greuter, and G.H. de Bock (2022) Mammographic breast density and the risk of breast cancer: a systematic review and meta-analysis. The Breast 66. Cited by: §1, §3.1. [5] F. Bray, M. Laversanne, H. Sung, J. Ferlay, R. L. Siegel, I. Soerjomataram, and A. Jemal (2024) Global cancer statistics 2022: globocan estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: A Cancer Journal for Clinicians 74 (3). Cited by: §1. [6] H. Chen and A. L. Martel (2025) Breast cancer detection from multi-view screening mammograms with visual prompt tuning. In Artificial Intelligence and Imaging for Diagnostic and Treatment Challenges in Breast Care. Deep-Breath., Cited by: §2. [7] C.J. D’Orsi, E.A. Sickles, E.B. Mendelson, and E.A. Morris (2014) ACR BI-RADS atlas: breast imaging reporting and data system. American College of Radiology. External Links: ISBN 9781559030168 Cited by: §3. [8] S. Dadsetan, D. Arefan, W. A. Berg, M. L. Zuley, J. H. Sumkin, and S. Wu (2022) Deep learning of longitudinal mammogram examinations for breast cancer risk prediction. Pattern Recognition 132. Cited by: §1. [9] K. Dembrower, P. Lindholm, and F. Strand (2020) A multi-million mammography image dataset and population-based screening cohort for the training and evaluation of deep neural networks—the cohort of screen-aged women (CSAW). Journal of Digital Imaging 33 (2). Cited by: §3. [10] J.A. Hanley and B.J. McNeil (1982) The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology 143. Cited by: §3. [11] K. He, X. Zhang, S. Ren, and J. Sun (2016) Deep residual learning for image recognition. In CVPR, Vol. . Cited by: §3. [12] G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger (2016) Deep networks with stochastic depth. In ECCV, Cited by: §2. [13] S. Hussain, M. Ali, U. Naseem, F. Nezhadmoghadam, M.A. Jatoi, T.A. Gulliver, and J.G. Tamez-Peña (2024) Breast cancer risk prediction using machine learning: a systematic review. Frontiers in Oncology 14. Cited by: §1. [14] J. J. Jeong, B. L. Vey, A. Bhimireddy, T. Kim, T. Santos, R. Correa, R. Dutt, M. Mosunjac, G. Oprea-Ilies, G. Smith, M. Woo, C. R. McAdams, M. S. Newell, I. Banerjee, J. Gichoya, and H. Trivedi (2023) The EMory BrEast imaging Dataset (EMBED): a racially diverse, granular dataset of 3.4 million screening and diagnostic mammographic images. Radiol Artif Intell 5 (1). Cited by: §3. [15] S. Jiang, D. L. Bennett, B. A. Rosner, and G. A. Colditz (2023) Longitudinal analysis of change in mammographic density in each breast and its association with breast cancer risk. JAMA Oncology 9 (6). Cited by: §1. [16] B. K. Karaman, K. Dodelzon, G. B. Akar, and M. R. Sabuncu (2024) Longitudinal mammogram risk prediction. In MICCAI, Cited by: §1, §1, §3, §3. [17] B. M. Keller, D. L. Nathan, Y. Wang, Y. Zheng, J. C. Gee, E. F. Conant, and D. Kontos (2012) Estimation of breast percent density in raw and processed full field digital mammography images via adaptive fuzzy c-means clustering and support vector machine segmentation. Medical Physics 39 (8). Cited by: §3. [18] H. Lee, J. Kim, E. Park, M. Kim, T. Kim, and T. Kooi (2023) Enhancing breast cancer risk prediction by incorporating prior images. In MICCAI, Cited by: §1, §2, §3. [19] I. Loshchilov and F. Hutter (2019) Decoupled weight decay regularization. Cited by: §3. [20] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala (2019) PyTorch: an imperative style, high-performance deep learning library. External Links: 1912.01703 Cited by: §3. [21] I.T. Rubio, C.A. Drukker, and A. Esgueva (2024) Risk-based breast cancer screening: what are the challenges?. Tumori J.. Cited by: §1. [22] Z. Sun, S. Thrun, and M. Kampffmeyer (2025) VMRA-MaR: An Asymmetry-Aware Temporal Framework for Longitudinal Breast Cancer Risk Prediction. In MICCAI, Cited by: §1, §1, §3.1, Table 2, Table 2, §3, §3. [23] S. Thrun, S. Hansen, Z. Sun, N. Blum, S. A. Salahuddin, X. Wang, K. Wickstrøm, E. Wetzer, R. Jenssen, M. Stille, and M. Kampffmeyer (2025) The impact of longitudinal mammogram alignment on breast cancer risk assessment. arXiv preprint arXiv:2511.08328. Cited by: §1, §2, §3.1, Table 2, Table 2, §3, §3. [24] S. Thrun, S. Hansen, Z. Sun, N. Blum, S. A. Salahuddin, K. Wickstrøm, E. Wetzer, R. Jenssen, M. Stille, and M. Kampffmeyer (2025) Reconsidering explicit longitudinal mammography alignment for enhanced breast cancer risk prediction. In MICCAI, Cited by: §1, §2, §3, §3. [25] H. Uno, T. Cai, M.J. Pencina, R.B. D’Agostino, and L.J. Wei (2011) On the c-statistics for evaluating overall adequacy of risk prediction procedures with censored survival data. Statistics in Medicine 30 (10). Cited by: §3. [26] C. Varela, N. Karssemeijer, J. H.C.L. Hendriks, and R. Holland (2005) Use of prior mammograms in the classification of benign and malignant masses. European Journal of Radiology 56 (2). Cited by: §1. [27] A.T. Wang, C.M. Vachon, K.R. Brandt, and K. Ghosh (2014) Breast density and breast cancer risk: a practical review. Mayo Clin. Proc. 89 (4). Cited by: §1, §3.1. [28] X. Wang, T. Tan, Y. Gao, E. Marcus, L. Han, A. Portaluri, T. Zhang, C. Lu, X. Liang, R. Beets-Tan, J. Teuwen, and R. Mann (2024) Ordinal Learning: Longitudinal Attention Alignment Model for Predicting Time to Future Breast Cancer Events from Mammograms . In MICCAI, Cited by: §1, §1, §2, §2, §3.1, Table 2, Table 2, §3, §3. [29] A. Yala, P. G. Mikhael, F. Strand, G. Lin, K. Smith, Y. Wan, L. Lamb, K. Hughes, C. Lehman, and R. Barzilay (2021) Toward robust mammography-based models for breast cancer risk. Sci. Transl. Med. 13 (578). Cited by: §1, §2, §2, §3.1, Table 2, Table 2, Table 2, Table 2, §3, §3. [30] N. Zielonke, A. Gini, E.E.L. Jansen, A. Anttila, N. Segnan, A. Ponti, P. Veerus, H.J. de Koning, N.T. van Ravesteyn, E.A.M. Heijnsdijk, S. Heinävaara, T. Sarkeala, M. Cañada, J. Pitter, G. Széles, Z. Voko, S. Minozzi, C. Senore, M. van Ballegooijen, I. Driesprong-de Kok, I. Lansdorp-Vogelaar, U. Ivanus, K. Jarm, D. Novak Mlakar, M. Primic-Žakelj, M. McKee, and J. Priaulx (2020) Evidence for reducing cancer-specific mortality due to screening for breast cancer in europe: a systematic review. Eur. J. Cancer 127. Cited by: §1.