Paper deep dive
REVEAL: Multimodal Vision-Language Alignment of Retinal Morphometry and Clinical Risks for Incident AD and Dementia Prediction
Seowung Leem, Lin Gu, Chenyu You, Kuang Gong, Ruogu Fang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 100%
Last extracted: 4/26/2026, 10:07:18 PM
Summary
REVEAL (REtinal-risk Vision-Language Early Alzheimer's Learning) is a multimodal framework designed to predict incident Alzheimer's Disease (AD) and dementia by aligning color fundus photographs with individualized clinical risk profiles. The framework addresses the modality gap between structured tabular data and natural language by using Llama 3.1 to generate synthetic clinical narratives. It introduces Group-Aware Contrastive Learning (GACL) to cluster patients with similar retinal morphometry and clinical characteristics, enhancing multimodal alignment. The model demonstrates the ability to predict disease onset an average of 8 years in advance, outperforming state-of-the-art retinal imaging models and general-purpose vision-language models.
Entities (10)
Relation Signals (7)
REVEAL ā aligns ā Color Fundus Photography
confidence 100% Ā· aligns color fundus photographs with individualized disease-specific risk profiles
Llama-3.1 ā generates ā Clinical Narratives
confidence 100% Ā· Using the LLaMA-3.1 API as the text generation engine... we converted each participant's risk factor profile into a synthetic clinical report.
RETFound ā isimageencoderfor ā REVEAL
confidence 100% Ā· we used RETFound (Zhou et al., 2023) as the image encoder
GatorTron ā istextencoderfor ā REVEAL
confidence 100% Ā· and GatorTron (Yang et al., 2022) as the text encoder
REVEAL ā predicts ā Alzheimer's disease
confidence 100% Ā· predicting incident AD and dementia on average 8 years before diagnosis
UK Biobank ā providesdatafor ā REVEAL
confidence 100% Ā· Color fundus photographs (CFPs) and AD and dementia-related risk factors were obtained from the UK Biobank
REVEAL ā uses ā Group-Aware Contrastive Learning
confidence 100% Ā· We further propose a group-aware contrastive learning (GACL) strategy
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The retina provides a unique, noninvasive window into Alzheimer's disease (AD) and dementia, capturing early structural changes through morphometric features, while systemic and lifestyle risk factors reflect well-established contributors to disease susceptibility long before clinical symptom onset. However, current retinal analysis frameworks typically model imaging and risk factors separately, limiting their ability to capture joint multimodal patterns critical for early risk prediction. Moreover, existing methods rarely incorporate mechanisms to organize or align patients with similar retinal and clinical characteristics, constraining the learning of coherent cross-modal associations. To address these limitations, we introduce REVEAL (REtinal-risk Vision-Language Early Alzheimer's Learning), a framework that aligns color fundus photographs with individualized disease-specific risk profiles for predicting incident AD and dementia, on average 8 years before diagnosis (range: 1-11 years). Because real-world risk factors are structured questionnaire data, we translate them into clinically interpretable narratives compatible with pretrained vision-language models (VLMs). We further propose a group-aware contrastive learning (GACL) strategy that clusters patients with similar retinal morphometry and risk factors as positive pairs, strengthening multimodal alignment. This unified representation learning framework substantially outperforms state-of-the-art retinal imaging models paired with clinical text encoders, as well as general-purpose VLMs, demonstrating the value of jointly modeling retinal biomarkers and clinical risk factors. By providing a generalizable and noninvasive approach for early AD and dementia risk stratification, REVEAL has the potential to enable earlier intervention and improve preventive care at the population level.
Tags
Links
- Source: https://arxiv.org/abs/2604.18757v1
- Canonical: https://arxiv.org/abs/2604.18757v1
Trouble viewing inline? Open PDF directly ā
Full Text
57,732 characters extracted from source content.
Expand or collapse full text
Proceedings of Machine Learning Research ā 139:1ā21, 2026Full Paper ā MIDL 2026 submission REVEAL: Multimodal VisionāLanguage Alignment of Retinal Morphometry and Clinical Risks for Incident AD and Dementia Prediction Seowung Leem 1 leem.s@ufl.edu Lin Gu 2 rin.tani.e8@tohoku.ac.jp Chenyu You 3,4 chenyu.you@stonybrook.edu Kuang Gong 1 KGong@bme.ufl.edu Ruogu Fang ā1 Ruogu.Fang@bme.ufl.edu 1 J. Crayton Pruitt Family Department of Biomedical Engineering, University of Florida, United States 2 Research Institute of Electrical Communication, Tohoku University, Japan 3 Department of Applied Mathematics & Statistics, Stony Brook University, United States 4 Department of Computer Science, Stony Brook University, United States Editors: Accepted for publication a MIDL 2026 Abstract The retina provides a unique, noninvasive window into Alzheimerās disease and demen- tia, capturing early structural changes through morphometric features, while systemic and lifestyle risk factors reflect well-established contributors to AD and dementia susceptibility long before clinical symptom onset. However, current retinal analysis frameworks typi- cally model imaging and risk factors separately, preventing them from capturing the joint multimodal patterns that are critical for early risk prediction. Moreover, existing methods rarely incorporate mechanisms to organize or align patients with similar retinal and clinical characteristics, limiting their ability to learn coherent cross-modal associations. To address these limitations, we introduce REVEAL (REtinal-risk Vision-language Early Alzheimerās Learning) that aligns color fundus photographs with individualized disease-specific risk pro- files for incident AD and dementia prediction on average 8 years before diagnosis (range: 1ā11 years). Because real-world risk factors are structured questionnaire data, we first translate them into clinically interpretable narratives compatible with pretrained vision- language models (VLMs). We further propose a group-aware contrastive learning (GACL) strategy that clusters patients with similar retinal morphometry and risk factors as positive pairs, strengthening multimodal alignment. This unified representation-learning framework substantially outperforms state-of-the-art retinal imaging models paired with clinical text encoders, as well as general VLMs, demonstrating the value of jointly modeling retinal biomarkers and clinical risk factors. By providing a generalizable, noninvasive approach for early AD and dementia risk stratification, REVEAL has the potential to enable earlier interventions and improve preventive care at the population level. Keywords: Retinal morphometry, risk factors, Alzheimerās disease and related dementia, Vision-language alignment, Contrastive learning ā Corresponding Author Ā© 2026 C-BY 4.0, S. Leem, L. Gu, C. You, K. Gong & R. Fang. arXiv:2604.18757v1 [cs.CV] 20 Apr 2026 Leem Gu You Gong Fang F 1 F 2 F 3 R 1 R 2 R 3 Original Positive Pair Revised Positive Pair Negative Pair Retinal Morphology Extraction Clinical Report Generation Intra-Modality Pair Matrix for Contrastive Learning Image&TextLatentEncoding Figure 1: Schematic of clinical scenario and proposed method. 1. Introduction Alzheimerās disease and dementia are progressive neurodegenerative diseases that manifest years before clinical symptom onset. Early identification of individuals at risk is critical for timely intervention and prevention. The retina offers a unique, noninvasive window into AD and dementia. Retinal morphometric features, referring to a set of quantitative measurements characterizing the size, shape, and structure of retinal components, have been shown to reflect early neurodegenerative changes and amyloid-β or tau deposition in the brain (Cheung et al., 2021; Koronyo et al., 2017; Ravichandran et al., 2025; Byun et al., 2021; Snyder et al., 2016). Parallel to retinal alterations, AD and dementia risk is strongly influenced by systemic and lifestyle factors (Leshner et al., 2017; Sprecher et al., 2017; Xiong et al., 2023; Hayden et al., 2024; Husz Ģar et al., 2024; Livingston et al., 2024). While retinal morphometry captures early neurodegenerative signatures, risk factors provide complementary information on modifiable factors that contribute to disease susceptibility. This convergence suggests that jointly modeling retinal biomarkers and systemic risk factors could improve early AD and dementia prediction beyond what either modality can achieve alone. Despite this potential, current approaches typically analyze retinal images and risk fac- tors separately, limiting their ability to capture the complex multimodal relationships un- derlying preclinical AD and dementia. Conventional contrastive learning frameworks often fail to align patients who share both retinal and systemic risk characteristics, leading to over- looked clinical commonalities (Figure 1). Moreover, structured risk-factor data from ques- tionnaires cannot be directly incorporated into standard vision-language models (VLMs), which are pretrained on natural language, creating a modality gap. 2 REVEAL To address these challenges, we introduce REVEAL (REtinal-risk Vision-language Early Alzheimerās Learning), a novel VLM-based framework that integrates retinal mor- phometric features with individualized disease-specific risk profiles. Structured risk factors are first transformed into clinically meaningful narratives using large language models, en- abling seamless multimodal representation learning. We further propose a group-aware contrastive learning strategy that leverages intra-modality similarity to identify clini- cally aligned individuals, capturing shared pathophysiological patterns across subjects. This approach allows REVEAL to learn unified representations that more accurately reflect the interplay between retinal biomarkers and systemic risk factors, offering improved early AD and dementia risk stratification. Our work has the following contributions: ⢠We introduce REVEAL, the first framework to jointly model fundus images and individualized AD and dementia risk factors by translating structured questionnaires into clinically meaningful narratives compatible with pretrained VLMs. ⢠We propose a group-aware contrastive learning strategy that identifies subjects sharing similar retinal morphometry and risk profiles, enabling coherent and clinically aligned multimodal representation learning. ⢠REVEAL achieves state-of-the-art performance in predicting incident AD and demen- tia on average 8 years before clinical onset (AD: mean = 8.68 years, range = 2.38ā 11.58 years; dementia: mean = 8.49 years, range = 1.50ā11.58 years) over retinal-only, clinical-text, and general VLM baselines, providing a generalizable, noninvasive ap- proach for population-level early AD and dementia risk stratification. 2. Method 2.1. Overview of REVEAL Framework The REVEAL framework was designed to operate in two stages. First, it aligned fundus images with individualized AD and dementia risk factors using a CLIP-style contrastive learning approach with our novel image-text pairing strategy. This enabled the model to learn multimodal relationships between colored fundus photography (CFP) and biological, phenotypic, and clinical markers of preclinical AD and dementia. Second, the learned joint representations were utilized in a downstream classifier to predict incident, preclinical AD and dementia (see Section 3.1 for details). 2.2. Constructing Clinical Report and Group-Aware Labels for Contrastive Learning 2.2.1. Synthetic Clinical Report Generation Direct application of CLIP was not feasible because the risk factors are represented as structured, tabular variables rather than natural-language descriptions. However, align- ment of fundus images with risk factors required a shared representation space that VLM can operate on. To bridge the modality gap between structured risk-factor variables and the natural-language input required by VLMs, we synthesized standardized clinical-style narratives from tabular health data (Figure 2). This transformation enabled the VLM to interpret the tabular risk factors in a linguistically contextualized form and facilitates multimodal alignment between fundus images and clinical attributes relevant to AD and 3 Leem Gu You Gong Fang The subject is 30 yearsold white female. The subject has 1.3HDL, 27 BMI, 120systolic blood pressure... Synthetic Clinical Report Age: 30 Sex: Female Ethn: White BMI: 27 HDL: 1.3 SBP: 120 Depr: Yes Insomnia: Yes History: N/A Walk: Yes Water: 3 Cup Coffee: Daily DataTemplate Prompt Here is given data: Data Use this as a template to answer the question:Template. If the data is missing, please say that data is not available. Do not try to infer anything. Unless the data is missing, include all information in the given data. Question: Using the template, can you create a clinical report summary of the given data? Llama 3.1 Subject The subject is years old . The subject has HbA1C, HDL, BMI.... Figure 2: Schematic overview of how a synthetic clinical report is generated. dementia. Using the LLaMA-3.1 API as the text generation engine (Grattafiori et al., 2024), we converted each participantās risk factor profile into a synthetic clinical report. For each subject, the LLM was provided with (1) a template prompt, (2) the subjectās structured risk factor values, and (3) explicit instructions for generating a concise medical summary. The template was adapted from the āPatient Informationā section of the CARE clinical case report guideline (Gagnier et al., 2013), ensuring that the synthesized narratives follow established clinical documentation conventions. The input prompt was designed to map the tabular information 1:1 into a template to prevent potential variability (Appendix A). This process produced consistent, clinically meaningful text representations that enable seamless integration of structured health information into our multimodal predictive framework. 2.2.2. Group-aware Contrastive Learning Strategy Conventional CLIP-style frameworks often fail in the medical domain (Radford et al., 2021). Prior studies showed that naive CLIP approaches struggle to capture the complex seman- tic relationships between images and disease-level information, highlighting the need for domain-specific strategies (Wang et al., 2022; Eslami et al., 2023). In our context, individu- als sharing both retinal and systemic AD and dementia risk characteristics must be grouped during training, since conventional contrastive learning only treats image-text pairs from the same subject as positive matches. To mitigate these gaps and enable the model to capture shared pathophysiological patterns across different modalities, we designed a group-aware contrastive learning (GACL). To introduce explicit clinical grounding, our GACL leverages morphometric features extracted directly from CFP, rather than solely on latent represen- tations from image encoders. This addressed limitations from prior works that attempted to improve the shortcomings of conventional CLIP by introducing the image-level or latent- level similarity (Du et al., 2024; Wu et al., 2024), which lacked explicit clinical grounding to find phenomenologically similar individuals with clinical relevance. By this design, the REVEAL learns a contrastive objective that encourages the model to associate the patterns of retinal signals and risk profiles that are linked to those patterns. Thus, REVEAL learns a latent disease-risk manifold, not object-level semantics. The GACL was inspired by (Bulat et al., 2024). As shown in Figure 3, FāR NĆK and TāR NĆD denote the z-normalized morphometric feature matrix and l2-normalized embeddings clinical report matrix for all N samples in 4 REVEAL Image Encoder 1000 0110 0110 0001 T 1 T 2 T 3 T 4 T 1 T 2 T 3 T 4 Clinical Report Similarity Mask 1100 1100 0010 0001 F 1 F 2 F 3 F 4 F 1 F 2 F 3 F 4 Morphometry Similarity Mask Clinical Report Generation Llama 3.1 Trainable Frozen Image Projection Subject Information Text Projection Retinal Morphology Extraction Colored Fundus Photography 1100 1110 0110 0001 I 1 I 2 I 3 I 4 T 1 T 2 T 3 T 4 = Image Embedding Text Embedding ķ³ (ķ) ķ³ (ķ») ķ³ I 1 I 2 I 3 I 4 T 1 T 2 T 3 T 4 Group-Aware Image-Text Similarity Matrix Generated Report Group Similarity Mask Loss Function ķ ķ» 11 111 11 1 I 1 I 2 I 3 I 4 T 1 T 2 T 3 T 4 Final Label Matrix Indicator Function -1 -1 -1 -1 -1 -1 -1 -1 Text Encoder ķ³ (ķķķķķ) ā 1,ķķ ķæ (") ķæ $ =1 ā1,ķķ ķæ " ķæ $ =0 ā ā 1 0 Boolean: True Boolean: False Figure 3: Schematic overview of how GACL is performed. a training batch, respectively. Here, K represents the number of morphometric features and D denotes the dimensionality of the text embedding space. To quantify the pairwise relationship within each modality, we computed the intra-morphometry similarity matrices S (F) āR NĆN for the fundus image and intra-clinical report similarity matrix S (T) āR NĆN for text. S (F) = FĀ· F ⤠,and S (T) = TĀ· T ⤠.(1) Each entry in S (F) characterizes how similar the retinal morphometric profiles of two subjects are, with a larger value indicating closer structural resemblance. Likewise, S (T) captures the semantic similarity between the clinical report embeddings, demonstrating the degree to which two subjects share encoded risk factor profiles. To identify subjects with similar characteristics, we thresholded both similarity matrices using modality-specific thresholdsĻ F andĻ T , yielding binary similarity masks L (F) and L (T) . In each mask, a value of 1 (Boolean True) indicates a similar sample pair, while 0 (Boolean False) indicates a dissimilar pair. L (F) = ( 1(True),if S (F) >Ļ F 0(False), otherwise and L (T) = ( 1(True),if S (T) >Ļ T 0(False), otherwise (2) To integrate information across modalities, we obtained a group similarity mask L (group) by applying a logical OR operation between two modality-specific masks. Finally, the group similarity mask was mapped by an indicator function, resulting in a contrastive learning- compatible final label matrix L, where entries of 1 were preserved, and 0s were converted 5 Leem Gu You Gong Fang to -1. This formulation preserved similarity relationships across modalities, ensuring that image-text alignment benefits from both structural consistency (from morphometry) and semantic consistency (from clinical reports). By reinforcing agreement between intra-modal similarity, image-text pairings were improved to maximize the learning efficiency between retinal morphometric features and risk factors. L (group) = ( 1, if L (F) ⨠L (T) = 1, 0, otherwise and L = ( 1,if L (group) = 1, ā1, otherwise (3) 2.3. Image-Text Alignment Learning with REVEAL 2.3.1. REVEAL Architecture The REVEAL framework was built on a standard contrastive vision-language learning setup to capture joint patterns between fundus images and AD and dementia risk factors. As shown in Figure 3, we used RETFound (Zhou et al., 2023) as the image encoder and GatorTron (Yang et al., 2022) as the text encoder, adding only lightweight projection layers to align their feature dimensions. During each forward pass, a raw fundus image and its synthesized clinical report were encoded and projected into a shared latent space. Retinal morphometrics and clinical narratives were further integrated into the GACL procedure to construct a label matrix. Finally, a group-aware image-text similarity matrix was computed using image embedding, text embedding, and a label matrix. The trainable REVEAL components were denoted as āflameā in Figure 3. This design enabled REVEAL to leverage both retinal imaging priors from foundation models and semantic priors from clinically trained language models. 2.3.2. Contrastive learning With GACL, the conventional contrastive objective was no longer applicable because it accommodated only a single positive pair per sample. Therefore, we adopted the loss from the prior work (Bulat et al., 2024) to support multiple clinically aligned pairs. L =ā 1 N img N txt N img X i=1 N txt X j=1 log 1 1 + exp l ij (ās ij /Ļ + β) ! ,(4) N img and N txt denote the number of images and texts in a training batch. The label term l ij ā+1,ā1 is the (i,j)-th entry of the final label matrix L, with l ij = 1 indicating a similar (positive) imageātext pair and l ij = ā1 indicating a dissimilar (negative) pair. The similarity value s ij is computed as the cosine similarity between the corresponding i-th image and j-th text embeddings obtained from the REVEAL framework. The temperature parameter is fixed at Ļ = 0.07. The bias term β is introduced to stabilize early training by reducing the initial loss, which is otherwise dominated by the large number of negative pairs. Including β, all hyperparameters (learning rate, eps, weight decay) and similarity thresholds (Ļ F andĻ T ) were chosen using an Optuna, hyperparameter optimization framework (Akiba et al., 2019), which identifies the optimal configuration within user-defined search ranges (details in Appendix B). 6 REVEAL 2.4. Study Population and Data Preprocessing 2.4.1. Subject Selection Color fundus photographs (CFPs) and AD and dementia-related risk factors were obtained from the UK Biobank (Sudlow et al., 2015). A total of 39,242 participants with high- quality CFPs were included and allocated into training (n=30,462), validation (n=3,384), and test (n=5,396) sets (Table 1, preprocessing details in Section 2.4.2). These splits each served a distinct role within the REVEAL framework. The training and validation sets were used solely in Stage 1 for representation alignment, with the validation set guiding hyperparameter tuning and similarity-threshold selection, while the test set was reserved for Stage 2 AD and dementia prediction. All participants who later developed incident AD or dementia were assigned to the test set, and only participants free of both prevalent and incident disease were included in the training and validation sets. Incident diagnoses were identified using UK Biobank dementia fields (42018, 42020, 42022, 42024). Among individuals with high-quality CFPs, 86 developed incident AD (mean time to diagnosis: 8.68 years; range: 2.38ā11.58) and 93 developed dementia of any subtype (mean: 8.49 years; range: 1.50ā11.58). To form the final evaluation cohort, control subjects without incident AD and dementia were sampled from the test pool to achieve an approximate 12% disease prevalence, consis- tent with estimates for adults agedā„65 years (Xiaopeng et al., 2025), while maintaining age and gender matched distributions (AD controls = 1,077; dementia controls = 1,139). From this cohort, 931 subjects (862 controls, 69 AD) for AD prediction and 985 subjects (911 controls, 74 dementia) for dementia prediction were used to train SVM models with 5-fold cross-validation. The remaining subjects, 232 (215 controls, 17 AD) for AD prediction and 247 (228 controls, 19 dementia) for dementia prediction, were held out as an independent test set. Cohort characteristics for downstream prediction tasks are provided in Tables 6 and 7 (Appendix C), and the distribution of onset years for AD and dementia are shown in Figure 4 (Appendix D) 2.4.2. Risk Factor Compilation and Retinal Image Processing A comprehensive set of demographic, behavioral, cognitive, and lifestyle variables was com- piled as risk factors based on established epidemiological evidence (Leshner et al., 2017; Sprecher et al., 2017; Xiong et al., 2023; Hayden et al., 2024; Husz Ģar et al., 2024; Livingston et al., 2024). The full list of these risk factors are provided in Appendix E. For the CFPs, image preprocessing and retinal morphometric feature extraction were carried out using the AutoMorph fundus morphology quantification pipeline (Zhou et al., 2022). A total of 136,994 CFPs were available from the initial UK Biobank assessment visit. AutoMorph Table 1: Demographic characteristics of the UK Biobank participants across the training, validation, and test cohorts. Train (n=30,462) Validation (n=3,384) Test (n=5,396) Gender: (male %)45.1045.4145.10 Age: mean (s.d)55.53 (8.24)55.78 (8.12)55.52 (8.17) Ethnicity: (British %)84.0883.5188.51 7 Leem Gu You Gong Fang first applied a convolutional neural networkābased quality-control module that classified images as low, moderate, or good quality. Following automated quality filtering and sub- sequent manual review, 66,251 high-quality images from 39,242 participants were retained for analysis. From these curated images, AutoMorph produced a structured set of retinal morphometric features (K=17; full list provided in Appendix F). These structural features have been shown in prior research to exhibit measurable differences in both preclinical and clinical stages of AD and dementia (Frost et al., 2013; Sharafi et al., 2019; Valenti, 2011; Ong et al., 2014; Armstrong et al., 2021). To maintain consistent anatomical orientation across eyes, all right-eye images were horizontally flipped before feature extraction. 3. Experiments 3.1. Downstream Tasks 3.1.1. Incident AD and incident dementia prediction We evaluated REVEAL on two prediction tasks: incident AD and incident dementia. For both tasks, we trained a multimodal SVM with an RBF kernel to perform binary classifi- cation, distinguishing individuals who later developed AD and dementia (normal at initial baseline visit and diagnosis reported after 1-11 years after baseline) from those who re- mained cognitively normal. The SVM produced probabilistic outputs, providing likelihood estimates for being AD/dementia-positive versus control. Each subject was represented by a concatenated multimodal feature vector composed of L2-normalized CFP image embeddings and text embeddings extracted from the REVEAL encoders. Class-weighted training was used to mitigate the imbalance between incident cases and controls. SVM hyperparameters (C and γ) were tuned using 5-fold cross-validation, and the best-performing model was sub- sequently evaluated on the independent hold-out test set. All reported results correspond to this final evaluation. 3.1.2. Comparison models To evaluate REVEAL, we compared its performance with several strong fundus-based foun- dation models: RETFound (CFP) (Zhou et al., 2023), RET-CLIP (Du et al., 2024), and KeepFIT-CFP (Wu et al., 2024), as well as medical multimodal vision-language models trained on multiple medical imaging types, including PMC-CLIP (Lin et al., 2023) and BiomedCLIP (Zhang et al., 2025). Because RETFound was an image encoder-only model, we paired it with GatorTron (Yang et al., 2022) to enable both image and text representa- tion. In the analysis, embeddings from two models were simply concatenated. In addition to these baselines, we trained a tabular SVM using clinical variables and CFP-derived mor- phometric features, applying most-frequent imputation for categorical variables and median imputation for continuous variables. Specifically, we tested tabular risk factors and mor- phometric features and risk factors with CFP latents to evaluate whether the improvement stems from the semantic richness of the LLM narrative or simply the power of the image foundation model. All models followed the same training and testing protocol as the mul- timodal SVM. Each experiment was repeated 10 times with different random seeds, and we report the average performance across runs. We used Welchās t-test and Hedgeās g to evaluate the statistical difference between REVEAL and comparison methods. 3.1.3. Threshold Evaluation of REVEAL Framework In REVEAL, thresholdsĻ F , andĻ T from GACL determine which image-text pairs should be grouped to share information, to learn shared representations among phenomenologically 8 REVEAL similar samples. Thresholds that are too low introduce noise by aligning dissimilar pairs, whereas thresholds that are too high restrict the modelās ability to capture meaningful cross-modal relationships. To assess their influence on predictive performance, we trained the model using varying threshold configurations. In each experiment, one threshold was fixed at the optimal value determined during optimization, while the other was varied systematically. Threshold candidates were chosen from the quartiles of the morphometric and text similarity distributions in the development set. 3.1.4. Evaluating Clinically Grounded Similarity in GACL As previously noted in Section 2.2.2, prior works have attempted to remedy the shortcomings of conventional CLIP by incorporating image-level or latent-level similarity. To evaluate the contribution of clinically grounded similarity in GACL, we compared downstream prediction performance under two configurations: (1) GACL using morphometric features as the source of image-image similarity, and (2) GACL using similarity computed directly from the image embeddings produced by the image encoder. This comparison allowed us to isolate the benefit of explicit clinical grounding for identifying phenotypically similar subjects and enhancing downstream AD and dementia prediction. 3.1.5. Evaluating the Effect of Different Logical Operators in GACL In preclinical disease settings, phenotypic similarity across different modalities can emerge asynchronously. For instance, individuals may share clinical risk factors indicative of el- evated neurodegenerative disease risk, while corresponding retinal signatures may not yet be present. To account for this asynchrony, GACL adopted a logical OR operator when defining group-level similarity. To validate this design choice, we conducted a compara- tive analysis using the logical AND operator. Specifically, we replaced the OR operator in Equation 3 with an AND operator while keeping all other parameters fixed. 3.2. Result Table 2: Performance of the incident Alzheimerās Disease prediction task. The average of 10 random seeds is presented as mean±std. The best results for each modality are in bold text. See Table 8 for statistics and effect size. AUROCBalanced Accuracy F1-ScoreMCC Baseline SVM0.593±0.0680.574±0.0830.140±0.0890.076±0.099 KeepFIT-CFP0.503±0.0610.519±0.0410.117±0.0380.018±0.045 BiomedCLIP0.525±0.0660.522±0.0520.121±0.0550.023±0.057 RETCLIP0.558±0.0760.527±0.0420.106±0.0690.028±0.051 PMC-CLIP0.471±0.0520.484±0.0200.076±0.024-0.022±0.024 RETFound+GatorTron 0.655±0.0600.573±0.0570.174±0.0980.108±0.095 Ours (no GACL)0.654±0.0970.602±0.0780.205±0.1010.144±0.111 Ours (with GACL) 0.658±0.095 0.610±0.083 0.208±0.105 0.147±0.117 3.2.1. Group-aware contrastive learning improves the Incident AD and dementia Prediction In the incident AD prediction task (Table 2), REVEAL achieved the best performance across nearly all evaluation metrics, including AUROC, balanced accuracy, F1-Score, and 9 Leem Gu You Gong Fang Matthewās Correlation Coefficient (MCC). Notably, the multimodal SVM trained on RE- VEAL embeddings substantially outperformed a baseline SVM trained directly on tabular risk factors and raw retinal morphometric features, demonstrating that vision-language em- beddings effectively transform raw modalities into enriched representations. Incorporating GACL further improved performance by aligning patients with similar retinal morphometry and risk profiles, enhancing overall predictive power. In the broader incident-dementia pre- diction task (Table 3), the SVM using REVEAL embeddings again outperformed baseline SVMs and other vision-language models. These results indicate that group-aware alignment strengthens multimodal representation learning, in both AD and dementia cases, demon- strating that retinal structural features closely correspond to disease-specific biomarkers. Statistical analysis of AD and dementia (Tables 8 and 9 in Appendix G) shows that these improvements are highly significant and associated with large effect sizes when compared to conventional CLIP-based models and SVM baselines. While comparisons with RET- Found+GatorTron do not always reach conventional statistical significance, these tests are based on 10 independent runs and are therefore underpowered to detect small-to- moderate effects. Importantly, GACL consistently improves predictive performance with non-negligible effect sizes, indicating meaningful practical gains rather than equivalence. The consistent improvements introduced by GACL highlight its effectiveness in enhancing representation learning for long-term neurodegenerative disease risk prediction. Impor- tantly, all CFPs in embedding learning were collected from cognitively normal participants at baseline, emphasizing that REVEAL, combined with a multimodal SVM, can identify preclinical AD and dementia risk by leveraging the complementary information between retinal morphometry and systemic risk factors. We further conducted an ablation study to examine the contribution of individual com- ponents in our model for incident AD and dementia prediction (Table 10 in Appendix H). Specifically, we evaluated the model using image embeddings alone (Image-only), image embeddings combined with raw tabular risk factors (Image+Table), and text embeddings alone (Text-only). Across both prediction tasks, Text-only representations consistently out- performed both Image-only and Image+Table variants, suggesting that clinical narratives capture substantially richer signals relevant to neurodegenerative disease risk. Notably, the joint Image-Text representation (REVEAL) achieved the best overall performance across all evaluation metrics, indicating that the enriched image representations provide comple- mentary information beyond text alone. In contrast, the Image+Table configurations un- derperformed the Text-only model, despite incorporating structured clinical variables. This finding highlights the advantage of replacing raw tabular features with clinical narratives, underscoring the benefit of higher-level semantic abstractions over simple concatenation between different model features. 3.2.2. Impact of Thresholds on REVEAL Performance The relative percentage differences in downstream prediction performance between the model trained with the optimal threshold and those trained under varyingĻ F andĻ T in REVEAL are shown in Figure 5 of Appendix I. Compared to the performance metric from the REVEAL with the optimal threshold (gray horizontal line, where values below 0 indicate worse performance and values above 0 indicate improvement), other models trained with different image or text thresholds did not yield better performance in most cases for both AD and dementia. For AD, using the highestĻ F produced the best accuracy, F1-score, 10 REVEAL Table 3: Performance of the incident dementia prediction task. The average of 10 random seeds is presented as mean± standard deviation. The best results for each modality are in bold text. See Table 9 for statistics and effect size. AUROCBalanced Accuracy F1-ScoreMCC Baseline SVM0.571±0.0920.572±0.0410.151±0.0420.075±0.041 KeepFIT-CFP0.487±0.0380.505±0.0410.110±0.0320.005±0.040 BiomedCLIP0.487±0.0430.502±0.0270.079±0.046-0.002±0.037 RETCLIP0.538±0.0870.547±0.0330.130±0.0400.051±0.037 PMC-CLIP0.484±0.0480.474±0.0300.054±0.039-0.031±0.033 RETFound+GatorTron 0.640±0.0620.577±0.0670.183±0.0950.121±0.101 Ours (no GACL)0.653±0.0720.596±0.0700.187±0.0920.135±0.096 Ours (with GACL) 0.659±0.073 0.605±0.070 0.189±0.091 0.140±0.096 and MCC, but at the cost of a reduced AUROC. This highlights the importance of carefully calibrated thresholds, as multimodal associations are highly sensitive to pairing phenotyp- ically similar pairs and avoiding weakly related alignments. Distinct trends were observed between image and text modalities. For images, higher thresholds demonstrated better performance, suggesting that lower thresholds introduce noise by forcing dissimilar samples to be similar. Conversely, for text embeddings, lower thresholds led to higher predictive performance, indicating that learning benefits when a broader range of semantically related texts are considered similar. In addition, the observed trade-off between accuracy, F1-score, MCC, and AUROC at higher image thresholds in incident AD prediction reflects a point estimate classification performance and ranking-based discrimination in prediction perfor- mance analysis using AUROC. In the incident AD prediction task, a higher image threshold forced the stricter alignment, which improved classification performance at a fixed operating point. However, the ranking ability across different thresholds was reduced as a tradeoff, leading to a lower accuracy. Therefore, the generalizability of these trends requires fur- ther validation in other domains and different datasets to validate the threshold-dependent trade-off influenced by dataset-specific factors. 3.2.3. Impact of Clinical vs. Latent Similarity in GACL The AD and dementia prediction results using morphometric features and the modelās la- tent features in image-image similarity computation of GACL are shown in Table 4. For this experiment, the threshold for the image latent was determined as the third quartile of the similarity distribution in the development set (Ļ F =0.9974). In both the incident AD and dementia prediction cases, incorporating morphometric features consistently yielded superior performance. This indicates that clinically grounded morphometric similarity pro- vides a more reliable and meaningful signal for identifying individuals who share similar retinal and systemic phenotypes, enabling richer and more discriminative representational learning. 3.2.4. Impact of Different Logical Operators in GACL The comparative analysis result between logical OR and AND operators in GACL for both incident AD and dementia prediction task are shown in Table 11 (Appendix J). Across 11 Leem Gu You Gong Fang both tasks, the OR and AND operators yielded nearly identical AUROC values, indicating comparable performance across different thresholds of the SVM classifier. However, the OR operator consistently achieved higher balanced accuracy, F1-Score, and MCC compared to the AND operator. This trend indicates that the requirement for similarity from at least one modality is more effective than a stricter similarity criterion. Thus, the OR operator provides greater flexibility by capturing partially overlapping phenotypic signals, leading to improved classification performance. Table 4: Performance of the incident AD and dementia prediction with different image- image similarity methods AUROCBalanced Accuracy F1-ScoreMCC AD Latent Feature0.656±0.0620.592±0.0790.201±0.1050.140±0.111 Morphometric Feature 0.658±0.095 0.610±0.083 0.208±0.105 0.147±0.117 Dementia Latent Feature0.654±0.0550.594±0.0520.181±0.0670.134±0.075 Morphometric Feature 0.659±0.073 0.605±0.070 0.189±0.091 0.140±0.096 4. Conclusion In this paper, we present REVEAL, a multimodal VLM framework that improves embedding learning for incident AD and dementia prediction by explicitly aligning retinal morphomet- ric features with individualized risk factors. Our group-aware contrastive learning strategy identifies clinically meaningful groups and patients with similar retinal and risk profiles, and enhances cross-modal representation learning. This alignment improves AD and dementia prediction diagnosed after an average of 8 years after the baseline visit. These gains demon- strate that multimodal alignment reflects the strong correspondence between AD-specific risk factors and retinal structural features. Moreover, transforming structured clinical data into narrative form leverages the semantic richness of pretrained language models, further strengthening multimodal associations and boosting predictive performance. These results underscore the value of clinically contextualized representation learning in VLMs for early AD and dementia risk stratification. Despite promising results, several limitations should be acknowledged. First, the performance of the REVEAL is sensitive to the threshold selection in GACL, reflecting a trade-off between strict phenotypic alignment and preserving suffi- cient shared representation for robust learning. Second, our evaluation is limited to a single large cohort (UK Biobank) with a limited number of incident cases of AD and dementia, limiting the generalizability of REVEAL to other populations and other disease settings. Finally, the evaluation of prompt variants for better alignment performance should be fur- ther evaluated. While absolute predictive performance remains limited by cohort size and disease prevalence, the consistent relative gains demonstrate the value of clinically grounded multimodal alignment for long-horizon neurodegenerative risk modeling. 12 REVEAL Acknowledgments This research has been conducted using data from UK Biobank, a major biomedical database under application ID 48388. This material is based upon work supported by the National Science Foundation under Grant No. (NSF 2123809). References Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. Op- tuna. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 2623ā2631, New York, NY, USA, July 2019. ACM. Grayson W Armstrong, Leo A Kim, Filippos Vingopoulos, et al. Retinal imaging findings in carriers with PSEN1-associated early-onset familial alzheimer disease before onset of cognitive symptoms. JAMA Ophthalmol., 139(1):49ā56, January 2021. Adrian Bulat, Yassine Ouali, and Georgios Tzimiropoulos. F: Fixing flawed foundations in contrastive pre-training results in very strong vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14172ā 14182, 2024. Min Soo Byun, Sung Wook Park, Jun Ho Lee, Dahyun Yi, So Yeon Jeon, Hyo Jung Choi, Haejung Joung, Un Hyung Ghim, Un Chul Park, Yu Kyeong Kim, Seong A Shin, Hyeong Gon Yu, Dong Young Lee, and KBASE Research Group. Association of retinal changes with alzheimer disease neuroimaging biomarkers in cognitively normal individu- als. JAMA Ophthalmol., 139(5):548ā556, May 2021. Carol Y. Cheung, Vincent Mok, Paul J. Foster, Emanuele Trucco, Christopher Chen, and Tien Yin Wong. Retinal imaging in Alzheimerās disease. Journal of Neurology, Neu- rosurgery & Psychiatry, 92(9):983ā994, September 2021. ISSN 0022-3050, 1468-330X. doi: 10.1136/jnnp-2020-325347. URL https://jnnp.bmj.com/content/92/9/983. Pub- lisher: BMJ Publishing Group Ltd Section: General neurology. Jiawei Du, Jia Guo, Weihang Zhang, Shengzhu Yang, Hanruo Liu, Huiqi Li, and Ningli Wang. RET-CLIP: A Retinal Image Foundation Model Pre-trained with Clinical Diagnos- tic Reports, August 2024. URL http://arxiv.org/abs/2405.14137. arXiv:2405.14137 [cs]. Sedigheh Eslami, Christoph Meinel, and Gerard de Melo. PubMedCLIP: How Much Does CLIP Benefit Visual Question Answering in the Medical Domain?In Andreas Vlachos and Isabelle Augenstein, editors, Findings of the Association for Computa- tional Linguistics: EACL 2023, pages 1181ā1193, Dubrovnik, Croatia, May 2023. As- sociation for Computational Linguistics. doi: 10.18653/v1/2023.findings-eacl.88. URL https://aclanthology.org/2023.findings-eacl.88/. S Frost, Y Kanagasingam, H Sohrabi, J Vignarajan, P Bourgeat, O Salvado, V Villemagne, C C Rowe, S Lance Macaulay, C Szoeke, K A Ellis, D Ames, C L Masters, S Rainey- Smith, R N Martins, and AIBL Research Group. Retinal vascular biomarkers for early detection and monitoring of alzheimerās disease. Transl. Psychiatry, 3(2):e233, February 2013. 13 Leem Gu You Gong Fang Joel J Gagnier, Gunver Kienle, Douglas G Altman, David Moher, Harold Sox, David Riley, and CARE Group*. The CARE guidelines: Consensus-based clinical case reporting guideline development. Glob. Adv. Health Med., 2(5):38ā43, September 2013. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, et al. The Llama 3 Herd of Models, November 2024. URL http://arxiv.org/abs/2407.21783. arXiv:2407.21783 [cs]. K.M. Hayden, M.M. Mielke, J.K. Evans, R. Neiberg, D. Molina-Henry, M. Culkin, S. Mar- covina, K.C. Johnson, O.T. Carmichael, S.R. Rapp, B.C. Sachs, J. Ding, H. Shappell, L. Wagenknecht, J.A. Luchsinger, and M.A. Espeland. Association between Modifiable Risk Factors and Levels of Blood-Based Biomarkers of Alzheimerās and Related Dementias in the Look AHEAD Cohort. JAR life, 13:1ā21, January 2024. ISSN 2534-773X. doi: 10. 14283/jarlife.2024.1. URL https://pmc.ncbi.nlm.nih.gov/articles/PMC10775955/. Zsolt Husz Ģar, Alina Solomon, Marie Anne Engh, Vanda Koszov Ģacz, Tam Ģas Terebessy, Zsolt Moln Ģar, P Ģeter Hegyi, Andr Ģas Horv Ģath, Francesca Mangialasche, Miia Kivipelto, and G Ģabor Csukly. Association of modifiable risk factors with progression to demen- tia in relation to amyloid and tau pathology. Alzheimerās Research & Therapy, 16: 238, October 2024. ISSN 1758-9193. doi: 10.1186/s13195-024-01602-9. URL https: //pmc.ncbi.nlm.nih.gov/articles/PMC11515263/. Yosef Koronyo, David Biggs, Ernesto Barron, David S. Boyer, Joel A. Pearlman, William J. Au, Shawn J. Kile, Austin Blanco, Dieu-Trang Fuchs, Adeel Ashfaq, Sally Frautschy, Gregory M. Cole, Carol A. Miller, David R. Hinton, Steven R. Verdooner, Keith L. Black, and Maya Koronyo-Hamaoui. Retinal amyloid pathology and proof-of-concept imaging trial in Alzheimerās disease. JCI Insight, 2(16), August 2017. ISSN 0021-9738. doi: 10.1172/jci.insight.93621. URL https://insight.jci.org/articles/view/93621. Publisher: American Society for Clinical Investigation. Alan I. Leshner, Story Landis, Clare Stroud, and Autumn Downey, editors. Preventing Cognitive Decline and Dementia: A Way Forward. National Academies Press, Wash- ington, D.C., September 2017. ISBN 978-0-309-45959-4. doi: 10.17226/24782. URL https://w.nap.edu/catalog/24782. Weixiong Lin, Ziheng Zhao, Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. PMC-CLIP: Contrastive Language-Image Pre-training using Biomedical Doc- uments, March 2023. URL http://arxiv.org/abs/2303.07240. arXiv:2303.07240 [cs]. Gill Livingston, Jonathan Huntley, Kathy Y Liu, Sergi G Costafreda, Geir SelbƦk, Su- varna Alladi, David Ames, Sube Banerjee, Alistair Burns, Carol Brayne, Nick C Fox, Cleusa P Ferri, Laura N Gitlin, Robert Howard, Helen C Kales, Mika Kivim Ģaki, Eric B Larson, Noeline Nakasujja, Kenneth Rockwood, Quincy Samus, Kokoro Shirai, Archana Singh-Manoux, Lon S Schneider, Sebastian Walsh, Yao Yao, Andrew Sommerlad, and Naaheed Mukadam. Dementia prevention, intervention, and care: 2024 report of the Lancet standing Commission. The Lancet, 404(10452):572ā628, August 2024. ISSN 0140-6736. doi: 10.1016/S0140-6736(24)01296-0. URL https://doi.org/10.1016/ S0140-6736(24)01296-0. Publisher: Elsevier. 14 REVEAL Yi-Ting Ong, Saima Hilal, Carol Yim-Lui Cheung, et al. Retinal vascular fractals and cognitive impairment. Dement. Geriatr. Cogn. Dis. Extra, 4(2):305ā313, May 2014. Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Language Su- pervision, February 2021. URL http://arxiv.org/abs/2103.00020. arXiv:2103.00020 [cs]. Swetha Ravichandran, Peter J Snyder, Jessica Alber, Charles F Murchison, Lauren E Chaby, Andreas Jeromin, and Edmund Arthur. Association and multimodal model of retinal and blood-based biomarkers for detection of preclinical alzheimerās disease. Alzheimers. Res. Ther., 17(1):19, January 2025. Sayed Mehran Sharafi, Jean-Philippe Sylvestre, Claudia Chevrefils, Jean-Paul Soucy, Syl- vain Beaulieu, Tharick A Pascoal, Jean Daniel Arbour, Marc-Andr Ģe Rh Ģeaume, Alain Robillard, C Ģeline Chayer, Pedro Rosa-Neto, Sulantha S Mathotaarachchi, Ziad S Nasred- dine, Serge Gauthier, and Fr Ģed Ģeric Lesage. Vascular retinal biomarkers improves the de- tection of the likely cerebral amyloid status from hyperspectral retinal images. Alzheimers Dement. (N. Y.), 5(1):610ā617, October 2019. Peter J Snyder, Lenworth N Johnson, Yen Ying Lim, Cl Ģaudia Y Santos, Jessica Alber, Paul Maruff, and Brian Fern Ģandez. Nonvascular retinal imaging markers of preclinical alzheimerās disease. Alzheimers Dement. (Amst.), 4(1):169ā178, October 2016. Kate E. Sprecher, Rebecca L. Koscik, Cynthia M. Carlsson, Henrik Zetterberg, Kaj Blennow, Ozioma C. Okonkwo, Mark A. Sager, Sanjay Asthana, Sterling C. John- son, Ruth M. Benca, and Barbara B. Bendlin. Poor sleep is associated with CSF biomarkers of amyloid pathology in cognitively normal adults. Neurology, 89(5):445ā 453, August 2017.ISSN 0028-3878.doi: 10.1212/WNL.0000000000004171.URL https://pmc.ncbi.nlm.nih.gov/articles/PMC5539733/. Cathie Sudlow, John Gallacher, Naomi Allen, Valerie Beral, Paul Burton, John Danesh, Paul Downey, Paul Elliott, Jane Green, Martin Landray, Bette Liu, Paul Matthews, Giok Ong, Jill Pell, Alan Silman, Alan Young, Tim Sprosen, Tim Peakman, and Rory Collins. UK Biobank: An Open Access Resource for Identifying the Causes of a Wide Range of Complex Diseases of Middle and Old Age. PLoS Medicine, 12(3):e1001779, March 2015. ISSN 1549-1277. doi: 10.1371/journal.pmed.1001779. URL https://w. ncbi.nlm.nih.gov/pmc/articles/PMC4380465/. Denise A Valenti. Alzheimerās disease and glaucoma: imaging the biomarkers of neurode- generative disease. Int. J. Alzheimers. Dis., 2010:793931, January 2011. Zifeng Wang, Zhenbang Wu, Dinesh Agarwal, and Jimeng Sun. MedCLIP: Contrastive Learning from Unpaired Medical Images and Text. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Nat- ural Language Processing, pages 3876ā3887, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.emnlp-main.256. URL https://aclanthology.org/2022.emnlp-main.256/. 15 Leem Gu You Gong Fang Ruiqi Wu, Chenran Zhang, Jianle Zhang, Yi Zhou, Tao Zhou, and Huazhu Fu. M-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise, May 2024. URL http://arxiv.org/abs/2405.11793. arXiv:2405.11793 [cs]. Zhu Xiaopeng, Yu Jing, Lai Xia, Wang Xingsheng, Deng Juan, Long Yan, and Li Baoshan. Global burden of alzheimerās disease and other dementias in adults aged 65 years and older, 1991-2021: population-based study. Front. Public Health, 13:1585711, July 2025. Jiayue Xiong, Rozina Bhimani, and Lisa Carney-Anderson. Review of Risk Factors Asso- ciated With Biomarkers for Alzheimer Disease. The Journal of Neuroscience Nursing: Journal of the American Association of Neuroscience Nurses, 55(3):103ā109, June 2023. ISSN 1945-2810. doi: 10.1097/JNN.0000000000000705. Xi Yang, Aokun Chen, Nima PourNejatian, Hoo Chang Shin, Kaleb E. Smith, Christo- pher Parisien, Colin Compas, Cheryl Martin, Anthony B. Costa, Mona G. Flores, Ying Zhang, Tanja Magoc, Christopher A. Harle, Gloria Lipori, Duane A. Mitchell, William R. Hogan, Elizabeth A. Shenkman, Jiang Bian, and Yonghui Wu. A large language model for electronic health records. npj Digital Medicine, 5(1):1ā9, December 2022. ISSN 2398- 6352. doi: 10.1038/s41746-022-00742-2. URL https://w.nature.com/articles/ s41746-022-00742-2. Publisher: Nature Publishing Group. Sheng Zhang, Yanbo Xu, Naoto Usuyama, Hanwen Xu, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, Cliff Wong, Andrea Tupini, Yu Wang, Matt Mazzola, Swadheen Shukla, Lars Liden, Jianfeng Gao, Angela Crabtree, Brian Piening, Carlo Bifulco, Matthew P. Lungren, Tristan Naumann, Sheng Wang, and Hoi- fung Poon. BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs, January 2025. URL http://arxiv.org/abs/ 2303.00915. arXiv:2303.00915 [cs]. Yukun Zhou, Siegfried K. Wagner, Mark A. Chia, An Zhao, Peter Woodward-Court, Moucheng Xu, Robbert Struyven, Daniel C. Alexander, and Pearse A. Keane. Auto- Morph: Automated Retinal Vascular Morphology Quantification Via a Deep Learning Pipeline. Translational Vision Science & Technology, 11(7):12, July 2022. ISSN 2164- 2591. doi: 10.1167/tvst.11.7.12. URL https://doi.org/10.1167/tvst.11.7.12. Yukun Zhou, Mark A. Chia, Siegfried K. Wagner, Murat S. Ayhan, Dominic J. Williamson, Robbert R. Struyven, Timing Liu, Moucheng Xu, Mateo G. Lozano, Peter Woodward- Court, Yuka Kihara, Andre Altmann, Aaron Y. Lee, Eric J. Topol, Alastair K. Denniston, Daniel C. Alexander, and Pearse A. Keane. A foundation model for generalizable disease detection from retinal images. Nature, 622(7981):156ā163, October 2023. ISSN 1476- 4687. doi: 10.1038/s41586-023-06555-x. URL https://w.nature.com/articles/ s41586-023-06555-x. Publisher: Nature Publishing Group. Appendix A. Full template for clinical report generation Template: The subject is <age> years old <ethnic background> <sex>. The average total household of this subject is in between <economic status>. The subject has <HbA1C> HbA1C, <HDL> HDL, <BMI> BMI, <systolic blood pressure> systolic blood pressure, 16 REVEAL <diastolic blood pressure> diastolic blood pressure. For lifestyle, the subject is in <em- ployment status>. The subject is <smoking history>, has <depression>, has sleep depri- vation <sleep deprivation>, and drinks alcohol <alcohol use>. The subject had his first cannabis at age <age of cannabis initiation> and used cannabis <cannabis use> times. The subject visits family <frequency of family visit>, and <number of leisure activity>. For physical activity, the subject walks <duration of walked 10+ minutes> minutes <number of days/week of walked 10+ minutes> days per week, exercises moderately <duration of mod- erate activity> minutes for <number of days/week of moderate activity> days a week, and exercises vigorously <duration of vigorous exercise> minutes for <number of days/week of vigorous activity> days a week. For diet, the subject has <cooked vegetable intake> tablespoons of cooked vegetables, <raw vegetable intake> tablespoons of raw vegetables, <fresh fruit intake> tablespoons of fresh fruit, and <dried fruit intake> dried fruit. In addition, the subject has oily fish <oily fish intake>, non-oily fish <non oily fish intake>, processed meat <processed meat intake>, poultry <poultry intake>, beef <beef intake>, lamb <lamb intake>, and pork <pork intake>. The subject has <bread intake> slices of bread per week, with <spread type>. The subject drinks <milk type>, <tea intake> cups of tea, <coffee intake> cups of coffee, <water intake> cups of water per day. The subject puts <salt added to food> in his diet. For cognitive function, the subject remembered <numeric memory> digits in the numeric memory test, scored <fluid intelligence> in a fluid intelligence test, completed trail #1 in <trail-making test A duration> deciseconds with <trail-making test A error counts> errors, and completed trail #2 in <trail-making test B duration> deciseconds with <trail-making test B error counts> errors. When a risk factor was unavailable (e.g., age of cannabis initiation), the report stated: No cannabis use was reported at that age in the <age of cannabis initiation> section. Appendix B. Implementation details and hyperparameter discovery The dimension of the projection layer for both image and text encoders was fixed at 1024. The batch size was fixed at 128. The parameter search space and determined values for REVEAL are available in Table 5. The ranges forĻ F andĻ T were determined by the 3rd quartile to the 4th quartile range of retinal morphometric similarities and pseudo- clinical report similarity in 85% of the development set. Based on Optuna, learning rate was determined as 2.42e-4, eps was determined as 8.61e-7, weight decay was set to 0.0232, thresholds were determined asĻ F =0.9481 andĻ T =0.9808. When training without GACL, we used the standard InfoNCE loss. Table 5: Hyperparameter search space and optimal values HyperparameterRange (min, max)Optimal Value learning rate1e-6, 5e-42.42e-4 eps1e-9, 1e-68.61e-7 weight decay1e-6, 1e-10.0232 Ļ F 0.2853, 0.99490.9480 Ļ T 0.9548, 0.99790.9808 β-5, 0-0.6319 17 Leem Gu You Gong Fang Appendix C. Demographic information of incident AD and dementia subjects and controls Table 6: SVM train-test splits and demographic characteristics of subjects with incident Alzheimerās Disease (AD) and controls With incident AD (n=86) Without incident AD (n=1077) SVM train / SVM test 69/17862/215 Gender: # male (%)45 (52.33)550 (51.07) Age: mean (s.d)64.23 (3.81)64.31 (3.73) Ethnicity: caucasian %86.0597.55 Table 7: SVM train-test splits and demographic characteristics of subjects with incident dementia and controls With incident dementia (n=93) Without incident dementia (n=1139) SVM train /SVM test 74/19911/228 Gender: # male (%)50 (53.76)607 (53.29) Age: mean (s.d)64.54 (3.87)64.24 (3.84) Ethnicity: caucasian %86.0297.28 Appendix D. Distribution of disease onset of Alzheimerās Disease and Dementia Figure 4: The years until onset of Alzheimerās Disease and dementia. IQR denotes in- terquartile range. 18 REVEAL Appendix E. Full list of AD and dementia risk factors used in this study ⢠Demographic Information (d = 5): Age, sex, economic status, ethnic background, employment status ⢠General Health Information (d = 11): BMI, HbA1C, HDL, systolic/diastolic blood pressure, numeric memory, fluid intelligence, Trail-Making Test A/B duration and error counts ⢠Risk Factors (d = 6): Depression, sleep deprivation, alcohol use, smoking history, cannabis use, age of cannabis initiation ⢠Physical activity (d = 6): Number and Duration of days/week walked 10+ minutes, Number and Duration of days/week of moderate physical activity 10+ minutes, Num- ber and Duration of days/week of vigorous physical activity 10+ minutes ⢠Social and leisure activities (d = 2): Frequency of friend&family visit, number of leisure activity ⢠Dietary habits (d = 18): cooked vegetable intake, raw vegetable intake, fresh fruit intake, dried fruit intake, oily fish intake, non-oily fish intake, processed meat intake, poultry intake, beef intake, lamb intake, pork intake, milk type, spread type, bread intake, salt added to food, tea intake, coffee intake, water intake Appendix F. Full list of fundus-based retinal morphometry used in this study ⢠Optic nerve head features(k = 2): Vertical and horizontal cup-to-disc ratios. ⢠Vascular features (k = 15): Fractal dimension, fractal density, distance tortuosity, squared curvature tortuosity, and tortuosity density for artery, vein, and both com- bined. Appendix G. Statistical comparison of REVEAL with baseline and other multimodal methods for incident AD and Dementia prediction Table 8: Welchās t-test results and Hedgesā g effect sizes for model performance in incident AD prediction. Each cell reports the p-value and corresponding effect size. See Table 2 for absolute performance values AUROCBalanced Accuracy F1-ScoreMCC Baseline SVM0.09 (0.75)0.35 (0.41)0.13 (0.67)0.16 (0.63) KeepFIT-CFP0.00 (1.87)0.01 (1.32)0.03 (1.09)0.00 (1.39) BiomedCLIP0.00 (1.56)0.01 (1.21)0.04 (0.99)0.01 (1.29) RETCLIP0.02 (1.11)0.01 (1.20)0.02 (1.09)0.01 (1.26) PMC-CLIP0.00 (2.35)0.00 (1.98)0.00 (1.66)0.00 (1.92) RETFound+GatorTron 0.92 (0.04)0.26 (0.49)0.48 (0.31)0.42 (0.35) 19 Leem Gu You Gong Fang Table 9: Welchās t-test results and Hedgesā g effect sizes for model performance in incident dementia prediction. Each cell reports the p-value and corresponding effect size. See Table 3 for absolute performance values AUROCBalanced Accuracy F1-ScoreMCC Baseline SVM0.03 (1.01)0.22 (0.54)0.26 (0.50)0.08 (0.83) KeepFIT-CFP0.00 (2.82)0.00 (1.16)0.03 (1.09)0.00 (1.76) BiomedCLIP0.00 (2.73)0.00 (1.86)0.00 (1.45)0.00 (1.86) RETCLIP0.00 (1.43)0.03 (1.00)0.09 (0.80)0.02 (1.16) PMC-CLIP0.00 (2.70)0.00 (2.33)0.00 (1.84)0.00 (2.27) RETFound+GatorTron 0.53 (0.27)0.38 (0.38)0.89 (0.05)0.68 (0.18) Appendix H. Component-wise ablation results for REVEAL Table 10: Component-wise ablation results for REVEAL on incident Ad and dementia pre- diction. Image-only uses image embeddings alone; Image+Table combines image embed- dings with raw tabular risk factors; Text-only uses LLM-derived clinical narrative embed- dings; and Image+Text jointly models image and text embeddings. Modelās performance is reported as mean±standard deviation across 10 runs AUROCBalanced Accuracy F1-ScoreMCC AD Image-only0.561±0.0560.527±0.0390.117±0.0440.029±0.044 Image+Table0.587±0.0770.559±0.0750.131±0.0860.061±0.091 Text-only0.630±0.0740.573±0.0570.188±0.0990.111±0.105 Image+Text0.658±0.0950.610±0.0830.208±0.1050.147±0.117 Dementia Image-only0.518±0.0500.523±0.0370.116±0.0430.089±0.030 Image+Table0.559±0.0830.553±0.0560.134±0.0630.056±0.065 Text-only0.641±0.0420.583±0.0590.168±0.0760.105±0.086 Image+Text0.659±0.0730.605±0.0700.189±0.0910.140±0.096 20 REVEAL Appendix I. Impact of Thresholds on REVEAL Performance Alzheimerās (ķ ķ )Alzheimerās (ķ ķ¹ )Dementia (ķ ķ )Dementia (ķ ķ¹ ) Figure 5: Effect (% difference) of varying thresholds on the incident AD and dementia prediction task. Appendix J. Performance of REVEAL with OR and AND operation Table 11 compares logical OR and AND operations in the GACL. While both strategies yield comparable AUROC, the OR operation consistently achieves equal or slightly bet- ter Balanced Accuracy, F1-score, and MCC across both tasks, indicating that enforcing similarity in either modality is more effective than requiring simultaneous agreement in both. Table 11: Performance Comparison between OR and AND function in GACL AUROCBalanced Accuracy F1-ScoreMCC AD AND0.659±0.0940.607±0.0820.205±0.1030.144±0.115 OR0.658±0.0950.610±0.0830.208±0.1050.147±0.117 Dementia AND0.659±0.0750.602±0.0710.184±0.0900.135±0.095 OR0.659±0.0730.605±0.0700.189±0.0910.140±0.096 21