Paper deep dive
A Two-Stage Multi-Modal MRI Framework for Lifespan Brain Age Prediction
Dingyi Zhang, Ruiying Liu, Yun Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 99%
Last extracted: 4/27/2026, 12:52:23 AM
Summary
The paper proposes a two-stage multi-modal MRI framework for lifespan brain age prediction, covering the full spectrum from fetal to elderly stages. The architecture utilizes a two-stage process: first, a coarse-grained age stage classification using a Modality Mixture of Experts (MoE), and second, a fine-grained brain age regression using an Age Stage MoE. The model integrates T1-weighted (T1w), T2-weighted (T2w), and fractional anisotropy (FA) modalities via late fusion, allowing for robust performance even with missing modalities. Experimental results on in-domain and out-of-domain datasets demonstrate that the proposed method (Ours-M) outperforms existing baselines like SFCN, ORDER, and TSAN, showing significant improvements in Mean Absolute Error (MAE) and generalization.
Entities (17)
Relation Signals (6)
BCP → isindomainof → Lifespan Brain Age Prediction
confidence 100% · the in-domain datasets include BCP (Howell et al., 2019)
Two-Stage Multi-Modal MRI Framework → outperforms → SFCN
confidence 100% · our multi-modal model (Ours-M) achieves the best overall performance... outperforms its single-modality variant... and existing baselines.
Two-Stage Multi-Modal MRI Framework → outperforms → ORDER
confidence 100% · our multi-modal model (Ours-M) achieves the best overall performance... outperforms... ORDER
Two-Stage Multi-Modal MRI Framework → outperforms → TSAN
confidence 100% · our multi-modal model (Ours-M) achieves the best overall performance... outperforms... TSAN
Two-Stage Multi-Modal MRI Framework → uses → T1-weighted (T1w)
confidence 100% · integrates T1-weighted (T1w), T2-weighted (T2w), and fractional anisotropy (FA) modalities
Two-Stage Multi-Modal MRI Framework → uses → Mixture of Experts (MoE)
confidence 100% · with a Mixture of Experts (MoE) (Shazeer et al., 2017) architecture
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The accurate quantification of brain age from MRI has emerged as an important biomarker of brain health. However, existing approaches are often restricted to narrow age ranges and single-modality MRI data, limiting their capacity to capture the coordinated macro- and microstructural changes that unfold across the human lifespan. To address these limitations, we developed a multi-modal brain age framework to characterize the integrated evolution of brain morphology and white matter organization. Our model adopts a two-stage architecture, where modalities are processed independently and integrated via late fusion in both stages: first to classify each subject into one of six developmental stages, and then to estimate age within the predicted stage. This design enables a unified and lifespan-spanning assessment of brain maturity across diverse developmental periods.
Tags
Links
- Source: https://arxiv.org/abs/2604.16655v1
- Canonical: https://arxiv.org/abs/2604.16655v1
Trouble viewing inline? Open PDF directly →
Full Text
14,127 characters extracted from source content.
Expand or collapse full text
Medical Imaging with Deep Learning – Under Review 2026Short Paper – MIDL 2026 submission A Two-Stage Multi-Modal MRI Framework for Lifespan Brain Age Prediction Dingyi Zhang ∗1 dingyi.zhang@emory.edu Ruiying Liu ∗2 rliu60@emory.edu Yun Wang 1,2 yun.wang2@emory.edu 1 Department of Computer Science, Emory University, Atlanta, GA, USA 2 Department of Biomedical Informatics, Emory University, Atlanta, GA, USA Editors: Under Review for MIDL 2026 Abstract The accurate quantification of brain age from MRI has emerged as an important biomarker of brain health. However, existing approaches are often restricted to narrow age ranges and single-modality MRI data, limiting their capacity to capture the coordinated macro- and microstructural changes that unfold across the human lifespan. To address these limitations, we developed a multi-modal brain age framework to characterize the integrated evolution of brain morphology and white matter organization. Our model adopts a two- stage architecture, where modalities are processed independently and integrated via late fusion in both stages: first to classify each subject into one of six developmental stages, and then to estimate age within the predicted stage. This design enables a unified and lifespan-spanning assessment of brain maturity across diverse developmental periods. Keywords: Brain Age Prediction, Multi-Modal MRI, Lifespan Brain Development. 1. Introduction Brain age prediction aims to estimate the biological age of the brain from MRI scans, with the gap between predicted and chronological age serving as an indicator of neuroanatomical health (Kumari and Sundarrajan, 2024). With the rapid development of deep learning, a growing number of methods have been proposed for this task, demonstrating promising performance on various neuroimaging datasets (Baecker et al., 2021). However, existing methods still face three key limitations: a focus on restricted age ranges rather than the full lifespan, reliance on single-modality data that fails to capture complementary macro- and microstructural information, and limited robustness to incomplete multimodal inputs in real-world settings. To address these limitations, we propose a two-stage multi-modal framework for lifespan brain age prediction. Our main contributions are as follows: (1) we present the first brain age prediction framework spanning the complete human lifespan, from fetal to elderly stages, covering the full spectrum of brain development and aging; (2) we adopt a late fusion strat- egy that integrates T1-weighted (T1w), T2-weighted (T2w), and fractional anisotropy (FA) modalities without requiring multi-modal training, naturally handling missing modalities at inference; (3) we design a two-stage pipeline combining coarse-grained age stage classi- fication and fine-grained brain age regression, with a Mixture of Experts (MoE) (Shazeer et al., 2017) architecture to handle diverse age-specific and modality-specific patterns. ∗ Contributed equally © 2026 C-BY 4.0, D. Zhang, R. Liu & Y. Wang. arXiv:2604.16655v1 [eess.IV] 17 Apr 2026 Zhang Liu Wang Figure 1: Overview of the proposed two-stage framework for lifespan brain age prediction. Stage 1 predicts ˆs, and Stage 2 estimates ˆy via stage-conditioned regression. 2. Method 2.1. Problem Formulation Given a set of MRI scans x m m∈M from available modalities M⊆T1w, T2w, FA pre- processed via N4 bias correction, skull stripping (Hoopes et al., 2022), and 1 m isotropic resampling, our goal is to predict the brain age ˆy ∈R of a subject. To unify age repre- sentation across the full lifespan, fetal age w (gestational weeks) is mapped to y = w− 40, yielding a continuous age axis from prenatal to elderly stages. 2.2. Framework Age Stage Classification. In the first stage, we partition the human lifespan into six age stages: fetal (gestational weeks 20–40), neonatal (postnatal weeks 0–12), infant (months 3–24), child (years 2–17), adult (years 18–65), and elderly (years > 65). For each available modality m, a shared MAE (He et al., 2022) backbone encodes x m into a latent represen- tation, which is then processed by a Modality MoE with hard routing, where each expert specializes in a specific MRI modality. This produces a probability distribution p(s | x m ) over age stages, where s denotes the age stage. The predicted stage is obtained by aggre- gating all predictions: ˆs = arg max s X m p(s| x m ). Brain Age Regression. In the second stage, a regression network estimates brain age conditioned on ˆs from first stage, enabling fine-grained prediction within each stage. We introduce an Age Stage MoE, where each expert specializes in a specific lifespan stage to capture diverse macro- and microstructural patterns. The shared MAE backbone and Modality MoE are reused and fine-tuned to leverage representations learned during the first stage. Modalities are processed independently, and their predictions are aggregated via late fusion at inference to obtain ˆy, supporting missing modalities without retraining. 2 Multi-Modal Lifespan Age Prediction Table 1: Brain age estimation results (MAE / STD in years↓). Best values are bold. Method FetalNeonatalInfantChildAdultElderly In-domain SFCN0.87 / 0.100.62 / 0.040.45 / 0.300.96 / 1.242.90 / 3.936.54 / 6.96 ORDER0.21 / 0.070.05 / 0.050.32 / 0.211.28 / 1.153.82 / 3.828.92 / 8.14 TSAN0.14 / 0.400.12 / 0.160.65 / 1.241.97 / 2.453.71 / 4.836.30 / 8.36 Ours-S0.01 / 0.02 0.02 / 0.030.14 / 0.211.35 / 1.904.51 / 6.686.02 / 7.84 Ours-M0.01 / 0.02 0.02 / 0.02 0.11 / 0.151.17 / 1.633.97 / 5.865.77 / 7.31 Out-of-domain SFCN0.97 / 0.120.81 / 0.410.79 / 0.555.36 / 4.1929.00 / 8.2342.97 / 11.95 ORDER0.25 / 0.070.16 / 0.250.42 / 0.147.56 / 3.7830.07 / 12.4536.73 / 22.63 TSAN0.68 / 1.110.72 / 1.921.43 / 2.664.68 / 4.5745.71 / 11.6156.30 / 15.08 Ours-S0.04 / 0.070.67 / 5.060.16 / 0.146.26 / 11.72 12.73 / 12.436.14 / 6.89 Ours-M0.04 / 0.07 0.08 / 0.15 0.15 / 0.10 3.88 / 7.94 12.73 / 12.436.14 / 6.89 3. Experiments Experimental Setup. We evaluate our method on nine datasets, which are categorized into two groups: in-domain and out-of-domain. A subset of the in-domain datasets is used for training, while the out-of-domain datasets are used exclusively for testing. Both groups cover the full lifespan, ranging from fetal to elderly stages. Specifically, the in-domain datasets include BCP (Howell et al., 2019), dHCP (Eyre et al., 2021), HCP-A (Bookheimer et al., 2019), HCP-D (Somerville et al., 2018), and HCP-YA (Van Essen et al., 2013), while the out-of-domain datasets consist of ABCD (Casey et al., 2018), ADNI (Jack Jr et al., 2008), FeTA (Payette et al., 2023), and HBCD (Volkow et al., 2024). We compare our approach with several baseline methods, including ORDER (Shah et al., 2024), SFCN (Peng et al., 2021), TSAN (Cheng et al., 2021), as well as a single-modal variant of our model. Results. As shown in Table 1, our multi-modal model (Ours-M) achieves the best overall performance across both in-domain and out-of-domain datasets, demonstrating strong gen- eralization ability. It also consistently outperforms its single-modality variant, highlighting the benefit of integrating multiple MRI modalities. Notably, in early developmental stages (fetal and neonatal), the prediction error is close to zero, indicating that the proposed two- stage framework effectively captures age-specific characteristics in these challenging stages. 4. Conclusion We propose a two-stage multi-modal framework for lifespan brain age prediction from MRI, integrating T1w, T2w, and FA modalities via late fusion to accommodate missing modal- ities. Experiments on nine datasets spanning fetal to elderly stages demonstrate superior performance and strong generalization, with our method reducing MAE by 10% and 69% in in-domain and out-of-domain settings, respectively, compared to existing baselines. More- over, multi-modal integration yields further gains over the single-modal variant, with im- provements of 8% and 11%, respectively, confirming the importance of multi-modal design. 3 Zhang Liu Wang Acknowledgments This work was supported by NIH grants R00HD103912, R01MH133313 (Y.W.). References Lea Baecker, Rafael Garcia-Dias, Sandra Vieira, Cristina Scarpazza, and Andrea Mechelli. Machine learning for brain age prediction: Introduction to methods and clinical applica- tions. EBioMedicine, 72, 2021. Susan Y Bookheimer, David H Salat, Melissa Terpstra, Beau M Ances, Deanna M Barch, Randy L Buckner, Gregory C Burgess, Sandra W Curtiss, Mirella Diaz-Santos, Jen- nifer Stine Elam, et al. The lifespan human connectome project in aging: an overview. Neuroimage, 185:335–348, 2019. Betty Jo Casey, Tariq Cannonier, May I Conley, Alexandra O Cohen, Deanna M Barch, Mary M Heitzeg, Mary E Soules, Theresa Teslovich, Danielle V Dellarco, Hugh Garavan, et al. The adolescent brain cognitive development (abcd) study: imaging acquisition across 21 sites. Developmental cognitive neuroscience, 32:43–54, 2018. Jian Cheng, Ziyang Liu, Hao Guan, Zhenzhou Wu, Haogang Zhu, Jiyang Jiang, Wei Wen, Dacheng Tao, and Tao Liu. Brain age estimation from mri using cascade networks with ranking loss. IEEE Transactions on Medical Imaging, 40(12):3400–3412, 2021. Michael Eyre, Sean P Fitzgibbon, Judit Ciarrusta, Lucilio Cordero-Grande, Anthony N Price, Tanya Poppe, Andreas Schuh, Emer Hughes, Camilla O’keeffe, Jakki Brandon, et al. The developing human connectome project: typical and disrupted perinatal func- tional connectivity. Brain, 144(7):2199–2213, 2021. Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll ́ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16000–16009, 2022. Andrew Hoopes, Jocelyn S Mora, Adrian V Dalca, Bruce Fischl, and Malte Hoffmann. Synthstrip: skull-stripping for any brain image. NeuroImage, 260:119474, 2022. Brittany R Howell, Martin A Styner, Wei Gao, Pew-Thian Yap, Li Wang, Kristine Baluyot, Essa Yacoub, Geng Chen, Taylor Potts, Andrew Salzwedel, et al. The unc/umn baby connectome project (bcp): An overview of the study design and protocol development. NeuroImage, 185:891–905, 2019. Clifford R Jack Jr, Matt A Bernstein, Nick C Fox, Paul Thompson, Gene Alexander, Danielle Harvey, Bret Borowski, Paula J Britson, Jennifer L. Whitwell, Chadwick Ward, et al. The alzheimer’s disease neuroimaging initiative (adni): Mri methods. Journal of Magnetic Resonance Imaging: An Official Journal of the International Society for Magnetic Resonance in Medicine, 27(4):685–691, 2008. LK Soumya Kumari and R Sundarrajan. A review on brain age prediction models. Brain Research, 1823:148668, 2024. 4 Multi-Modal Lifespan Age Prediction Kelly Payette, Hongwei Bran Li, Priscille De Dumast, Roxane Licandro, Hui Ji, Md Mah- fuzur Rahman Siddiquee, Daguang Xu, Andriy Myronenko, Hao Liu, Yuchen Pei, et al. Fetal brain tissue annotation and segmentation challenge results. Medical image analysis, 88:102833, 2023. Han Peng, Weikang Gong, Christian F Beckmann, Andrea Vedaldi, and Stephen M Smith. Accurate brain age prediction with lightweight deep neural networks. Medical image analysis, 68:101871, 2021. Jay Shah, Md Mahfuzur Rahman Siddiquee, Yi Su, Teresa Wu, and Baoxin Li. Ordinal classification with distance regularization for robust brain age prediction. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 7882–7891, 2024. Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hin- ton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of- experts layer. arXiv preprint arXiv:1701.06538, 2017. Leah H Somerville, Susan Y Bookheimer, Randy L Buckner, Gregory C Burgess, Sandra W Curtiss, Mirella Dapretto, Jennifer Stine Elam, Michael S Gaffrey, Michael P Harms, Cynthia Hodge, et al. The lifespan human connectome project in development: A large- scale study of brain connectivity development in 5–21 year olds. Neuroimage, 183:456–468, 2018. David C Van Essen, Stephen M Smith, Deanna M Barch, Timothy EJ Behrens, Essa Yacoub, Kamil Ugurbil, Wu-Minn HCP Consortium, et al. The wu-minn human connectome project: an overview. Neuroimage, 80:62–79, 2013. Nora D Volkow, Joshua A Gordon, Diana W Bianchi, Michael F Chiang, Janine A Clayton, William M Klein, George F Koob, Walter J Koroshetz, Eliseo J P ́erez-Stable, Jane M Simoni, et al. The healthy brain and child development study (hbcd): Nih collaboration to understand the impacts of prenatal and early life experiences on brain development. Developmental Cognitive Neuroscience, 69:101423, 2024. 5 Zhang Liu Wang Appendix A. Dataset Distribution Figure 2: Dataset distribution. Top: age distribution across datasets. Bottom: number of samples per age group and dataset. Each subject-session counts as one sample. Figure 2 illustrates dataset distribution. The in-domain set comprises five datasets (BCP, dHCP, HCP-D, HCP-YA, and HCP-A) totaling 3804 subjects, 4,588 sessions and 11927 scans, while the out-of-domain set consists of four held-out test datasets (ABCD, ADNI, FeTA, and HBCD) with 1527 subjects, 1,675 sessions and 2251 scans in total. 6 Multi-Modal Lifespan Age Prediction Appendix B. Age Stage Classification Results Figure 3: Confusion matrices of the first stage classifier on the out-of-domain test set. Figure 3 shows the confusion matrices of the first stage on the out-of-domain test set. Note that the two matrices have different sample counts: Ours-S evaluates each modal- ity independently (2,688 samples), whereas Ours-M aggregates all available modalities per subject into a single prediction (1,675 subjects). Ours-M outperforms Ours-S in overall accuracy (88.18% vs. 80.62%) and eliminates extreme misclassifications: unlike Ours-S, which misclassifies some Neonatal subjects as Adult or Elderly, Ours-M confines all errors to neighboring age groups. Adult accuracy remains low for both models (29%), as out-of- domain adult subjects come exclusively from ADNI (ages 60–65), placing them near the Adult/Elderly boundary. 7