Paper deep dive
Gated Multi-Graph Fusion via Graph Attention Networks for Alzheimer's Disease Detection
Jinyu Li, Xiao Wei, Bin Wen, Kai Li, Yuqin Lin, Xiaobao Wang, Longbiao Wang, Jianwu Dang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 7/5/2026, 5:54:58 AM
Summary
The paper proposes a Multi-View Gated Graph Attention Network (GAT) for the non-invasive detection of Alzheimer's Disease (AD) through spontaneous speech analysis. The model utilizes a 'content-structure-flow' framework by constructing three distinct graphs: a Semantic Graph (content), a Dependency Graph (structure), and a PMI-based Co-occurrence Graph (flow) to capture narrative logic. An adaptive gated fusion mechanism is employed to address clinical heterogeneity by dynamically weighting these views. Evaluated on the ADReSSo 2021 dataset, the model achieved a 90.00% accuracy, outperforming several baseline methods and demonstrating the importance of the co-occurrence graph and gated fusion in capturing complex linguistic patterns.
Entities (9)
Relation Signals (7)
Multi-View Gated Graph Attention Network → detects → Alzheimer's disease
confidence 100% · We propose a Multi-View Gated Graph Attention Network... for Alzheimer's Disease Detection
BERT-base → generates → word-level embeddings
confidence 100% · To represent each word in a high-dimensional semantic space, we use a pre-trained BERT-base model.
Multi-View Gated Graph Attention Network → incorporates → Dependency Graph
confidence 100% · We propose a multi-view graph learning framework... Semantic, dependency, and co-occurrence graphs
Multi-View Gated Graph Attention Network → incorporates → Co-occurrence Graph
confidence 100% · We propose a multi-view graph learning framework... Semantic, dependency, and co-occurrence graphs
Multi-View Gated Graph Attention Network → incorporates → Semantic Graph
confidence 100% · We propose a multi-view graph learning framework... Semantic, dependency, and co-occurrence graphs
Whisper → performs → Automatic Speech Recognition (ASR)
confidence 100% · we utilize Whisper [18] for transcription.
Co-occurrence Graph → uses → Pointwise Mutual Information (PMI)
confidence 100% · the co-occurrence graph leverages Pointwise Mutual Information (PMI) from a normative corpus
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Spontaneous speech is a vital non-invasive biomarker for Alzheimer's Disease (AD), yet many systems overlook non-linear structural disruptions and clinical heterogeneity in pathological language. We propose a Multi-View Gated Graph Attention Network that transcribes audio via Automatic Speech Recognition (ASR) to construct semantic, dependency, and co-occurrence graphs, characterizing speech through a "content-structure-flow" framework. Notably, the co-occurrence graph leverages Pointwise Mutual Information (PMI) from a normative corpus to quantify narrative logic and linguistic deviation. To address symptomatic diversity, an adaptive gated fusion mechanism dynamically integrates these views. Evaluated on the ADReSSo dataset, our model achieves 90.00% accuracy. Ablation results confirm that the PMI-based graph and heterogeneity-aware gating are essential for robust classification across diverse clinical populations. Our source code is publicly available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2606.31186v1
- Canonical: https://arxiv.org/abs/2606.31186v1
Trouble viewing inline? Open PDF directly →
Full Text
28,710 characters extracted from source content.
Expand or collapse full text
Gated Multi-Graph Fusion via Graph Attention Networks for Alzheimer’s Disease Detection Jinyu Li 1 , Xiao Wei 1,2 , Bin Wen 1,2 , Kai Li 2 , Yuqin Lin 3 , Xiaobao Wang 1 , Longbiao Wang 1,4,∗ , Jianwu Dang 2,∗ 1 School of Future Technology, Tianjin University, Tianjin, China 2 Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China 3 College of Computer and Data Science, Fuzhou University, Fuzhou, China 4 Huiyan Technology (Tianjin) Co., Ltd, Tianjin, China lijinyu536@tju.edu.cn, longbiao wang@tju.edu.cn Abstract Spontaneous speech is a vital non-invasive biomarker for Alzheimer’s Disease (AD), yet many systems overlook non- linear structural disruptions and clinical heterogeneity in patho- logical language. We propose a Multi-View Gated Graph At- tention Network that transcribes audio via Automatic Speech Recognition (ASR) to construct semantic, dependency, and co-occurrence graphs, characterizing speech through a ”con- tent–structure–flow” framework. Notably, the co-occurrence graph leverages Pointwise Mutual Information (PMI) from a normative corpus to quantify narrative logic and linguistic de- viation. To address symptomatic diversity, an adaptive gated fusion mechanism dynamically integrates these views. Eval- uated on the ADReSSo dataset, our model achieves 90.00% accuracy. Ablation results confirm that the PMI-based graph and heterogeneity-aware gating are essential for robust classi- fication across diverse clinical populations. Our source code is publicly available at https://github.com/opeacc/AD. Index Terms: Alzheimer’s Disease Detection, Graph Neural Networks, Spontaneous Speech Analysis, Gated Fusion Mech- anism 1. Introduction Dementia, primarily AD, represents an escalating global health crisis hallmarked by progressive cognitive decline that mani- fests early in spontaneous speech [1, 2]. The ”Cookie Theft” picture description task is extensively recognized for its clini- cal utility in eliciting these critical linguistic markers in a con- trolled environment [3]. Initial efforts in speech-based dementia detection predominantly relied on traditional machine learning paradigms [4, 5]. These methodologies utilized manual feature engineering to extract handcrafted acoustic parameters—such as pitch and jitter—alongside lexical diversity metrics, which were subsequently classified using algorithms like Support Vec- tor Machines (SVM) or Random Forests [6, 7]. While these models offered high interpretability [8], they were constrained by their dependence on expert-defined features and their in- herent inability to capture the latent, non-linear dependencies found in spontaneous discourse. The subsequent evolution of deep learning, specifically the application of Large Language Models (LLMs) such as BERT, has substantially improved di- agnostic performance [9, 10]. Nonetheless, these models pri- marily learn semantic representations implicitly [11, 12, 13] and often fail to sufficiently characterize the complex, multi- * indicates the corresponding authors. dimensional structural degradation characteristic of patholog- ical speech. Recently, researchers have explored the use of Graph Neural Networks (GNNs) to model the intricate topolog- ical structure of language, yielding promising results [14, 15]. However, current research in this domain remains insufficient in its modeling of linguistic features. Furthermore, existing frameworks typically rely on simplistic fusion strategies (e.g., concatenation or fixed-weight summation), limiting the model’s adaptability to diverse symptomatic manifestations across the clinical spectrum. In this paper, we propose a multi-view graph learning framework that addresses two critical gaps in current re- search through the following innovations: First, we introduce a ”content–structure–flow” trinity for comprehensive linguis- tic modeling. Traditional methods often overlook the ”flow” of discourse—the logical progression of speech. While se- mantic graphs capture ”content” and dependency graphs cap- ture ”structure,” we specifically design a Co-occurrence Graph based on PMI derived from healthy normative data. This graph uniquely reflects a subject’s ability to describe local events. In patients with AD, speech often exhibits disordered sequences, logical jumps, and repetitive loops [16]. By integrating this co-occurrence perspective, our model provides a full-spectrum characterization of linguistic patterns, capturing not just what is said, but how the narrative logic unfolds. Second, we pro- pose a Gated Fusion Mechanism to address the clinical hetero- geneity of AD. Clinical observations demonstrate that dementia symptoms are not monolithic: some patients exhibit ”syntac- tic collapse” (simplified grammar with preserved vocabulary), while others suffer from ”semantic void” (fluent but meaning- less speech) [17]. Our gated network adaptively assigns weights to the semantic, dependency, and co-occurrence representations on a per-sample basis. This allows the framework to capture sample-specific features, dynamically focusing on the most dis- criminative linguistic markers for each individual. In summary, our contributions are as follows: 1. Discourse Flow Analysis: We utilize a PMI-based co- occurrence graph to quantify the deviation of event descrip- tion logic from healthy norms. 2. A Holistic Multi-Graph Framework: We integrate semantic, dependency, and co-occurrence graphs to model the ”con- tent–structure–flow” of spontaneous speech. 3. Heterogeneity-aware Fusion: We implemented a gated fu- sion mechanism that accounts for the diversity of AD symp- toms, enhancing the model’s ability to adapt to varying symp- tomatic presentations. arXiv:2606.31186v1 [cs.CL] 30 Jun 2026 Figure 1: Overview of the proposed Multi-View Gated Graph Attention Network framework for dementia detection. (1) Audio is transcribed via ASR and converted into word-level embeddings. (2) Semantic, co-occurrence, and dependency graphs are constructed to capture multi-scale linguistic patterns. (3) Topological features are extracted by GAT layers and aggregated via global pooling. (4) These multi-view features are adaptively fused by a gating network. (5) The fused features are fed into an MLP for final classification. 2. Multi-View Gated GAT We propose a multi-dimensional graph learning framework, as illustrated in Figure 1, designed to capture the ”con- tent–structure–flow” of spontaneous speech. This architecture consists of five main modules as detailed below. 2.1. Speech Transcription and Node Embedding Automatic Speech Recognition (ASR): Given the raw audio signal of a subject describing the ”Cookie Theft” picture, we utilize Whisper [18] for transcription. Whisper’s large-scale weak supervision training makes it exceptionally robust to the disfluencies common in dementia-impacted speech. The output is a sequence of N words: S =w 1 ,w 2 ,...,w N . Node Feature Initialization: To represent each word in a high-dimensional semantic space, we use a pre-trained BERT- base [19] model.Since BERT employs a WordPiece to- kenizer, a word w i may be split into k sub-word tokens t i,1 ,t i,2 ,...,t i,k . To maintain a 1-to-1 mapping between the transcript and graph nodes, we define the node representation x i ∈R 768 as the mean of its constituent token embeddings: x i = MeanPooling(BERT(t i,1 ),..., BERT(t i,k )).(1) The final node feature matrix for a transcript is X∈R N×768 . 2.2. Multi-View Graph Construction We define three graphsG sem , G syn , andG co sharing the same vertex setV , where|V| = N . 2.2.1. Semantic Graph: Content Representation This graph models the conceptual density and semantic rela- tionships. We compute the cosine similarity between every pair of word embeddings x i and x j [20]. An edge e ij ∈ E sem is created if the similarity exceeds a predefined threshold τ s : Sim(x i , x j ) = x ⊤ i x j ∥x i ∥x j ∥ > τ s .(2) This captures the global semantic structural characteristics of the subject. 2.2.2. Dependency Graph: Structural Integrity To capture grammatical decay (e.g., loss of complex clauses), we use spaCy [21] to perform dependency parsing. An edge e ij ∈ E syn is established if there is a direct syntactic depen- dency (e.g., nominal subject, direct object, modifier) between w i and w j . This graph reflects the structural complexity of the subject’s language [22, 23]. 2.2.3. Co-occurrence Graph: Discourse Flow via Normative PMI To capture the ”flow” dimension of the ”content-structure-flow” trinity, we design a co-occurrence graph that quantifies the logical progression and temporal coherence of speech. Tradi- tional sequential models often fail to detect the subtle ”logical jumps” or repetitive loops characteristic of AD-impacted dis- course. Our approach addresses this by measuring how much a subject’s word associations deviate from established healthy linguistic norms. We first construct a Normative CorpusD hc using only tran- scripts from healthy controls. This corpus serves as a linguistic reference baseline for characterizing natural word associations within the context of the ’Cookie Theft’ task. We calculate the Pointwise Mutual Information (PMI) [24, 25] for all word pairs (w i ,w j ) that appear within a sliding window of size n across D hc : PMI(w i ,w j ) = log P(w i ,w j ) P(w i )P(w j ) ,(3) where P(w i ) and P(w j ) are the individual unigram probabil- ities, and P(w i ,w j ) is the joint probability of co-occurrence within the window. A high PMI indicates a strong, statistically significant association between words (e.g., ”sink” and ”over- flowing”) that reflects healthy narrative logic. For a specific subject’s transcript S, we construct the co- occurrence graphG co by mapping the pre-calculated normative weights onto the subject’s actual word sequence. An edge e ij ∈ E co is established between word w i and any subsequent word w i+k (where 1 ≤ k ≤ n) if and only if their normative PMI exceeds a predefined threshold τ c . The structural density inherent in the co-occurrence graph G co functions as a direct proxy for discourse coherence, effec- tively capturing the nuanced temporal ”flow” of spontaneous speech. Within a healthy flow scenario—where a subject ad- heres to a logical and conventional narrative trajectory—the re- sultingG co is defined by a high density of edges paired with sig- nificant PMI weights. Together, these elements construct a ro- bustly well-connected topological ”skeleton” composed of nor- mative word associations. Conversely, the pathological disrup- tion observed in patients with AD—which manifests as labored word retrieval or disjointed logical transitions—produces word pairs that are statistically rare or nonsensical compared to the normative baseline. This ultimately yields a fragmented graph topology marked by sparse or anomalous connectivity. 2.3. View-Specific Encoding via GAT To learn the importance of different neighbors within each graph, we employ a Graph Attention Network (GAT) [26]. For a node i in graph k ∈sem,syn,co, the attention coefficient α ij is computed as: α ij = exp LeakyReLU a ⊤ Wx i ∥ Wx j P l∈N i exp LeakyReLU a ⊤ Wx i ∥ Wx l , (4) wherea and W are learnable parameters andN i is the neigh- borhood of node i. The updated node features h ′ i are aggregated to form a global graph representation z k using Global Mean Pooling: z k = 1 N N X i=1 h ′ i , z k ∈R d gat .(5) 2.4. Heterogeneity-Aware Gated Fusion Given that AD symptoms vary (e.g., some lose syntax, others lose semantic logic), we use a gated mechanism to dynamically weight the three views [27]. First, we concatenate the three graph vectors: Z cat = [z sem ∥ z syn ∥ z co ]. The gating network computes a weight vector g: g = Softmax(W g Z cat + b g ), g = [β sem ,β syn ,β co ]. (6) The fused representation z fused is a weighted sum: z fused = β sem z sem + β syn z syn + β co z co .(7) To prevent the loss of individual view details during fusion, we apply a pass-through concatenation: z final = z fused ∥ z sem ∥ z syn ∥ z co .(8) This ensures that the classifier has access to both the ”optimized mixture” and the raw specific features of each linguistic dimen- sion. 2.5. Classification and Objective Function The final vector z final is passed through a Multi-Layer Per- ceptron (MLP). To classify the subject as either AD or Healthy Control (HC), the model outputs a predicted probability ˆy = Sigmoid(MLP(z final )). To improve the model’s generaliza- tion and prevent over-confident predictions, we employ Label Smoothing during training. Instead of using hard binary labels y ∈0, 1, we transform them into soft targets y ls : y ls = y(1− α) + α K ,(9) where α = 0.2 is the smoothing parameter and K = 2 is the number of classes. The model is then optimized using the Smoothed Binary Cross-Entropy (BCE) Loss: L =− 1 M M X m=1 [y ls,m log(ˆy m )+(1−y ls,m ) log(1−ˆy m )]. (10) This approach encourages the model to learn more robust latent representations by penalizing extreme logit values, thereby en- hancing both calibration and classification performance on the clinical dataset. 3. Experiments 3.1. Experimental settings 3.1.1. Dataset Our experiments were conducted using the standardized ADReSSo 2021 Challenge dataset [28], which is a curated sub- set of the Pitt Corpus within the DementiaBank database. Sam- ples were collected via the ”Cookie Theft” picture description task—a clinical gold standard for assessing narrative speech and cognitive status. The corpus comprises audio recordings and transcripts from 237 English speakers, including 122 individu- als diagnosed with AD and 115 HC. Following the ADReSSo 2021 protocol, the data is partitioned into a training set of 166 participants (87 AD, 79 HC) and a test set of 71 participants (35 AD, 36 HC). We cropped the raw audio files based on provided timestamps to retain only segments containing subject speech. Since our approach constructs graph nodes directly from word- level tokens in subject narratives, samples lacking valid subject- specific segments—specifically five from the training set and one from the test set—were excluded as they could not support graph construction. To ensure experimental fairness and rigor- ous comparison, we re-implemented all baseline methods and evaluated them on this filtered version of the dataset. 3.1.2. Implementation details All experiments were conducted on an NVIDIA GeForce RTX 4090D GPU. For the model architecture, the input dimension was set to 768 to match the BERT-base embeddings. The GAT component consisted of a single layer with 2 attention heads and a hidden dimension of 128. The subsequent MLP utilized a hidden layer of 256 units. To prevent overfitting, we applied a dropout rate of 0.5 and a weight decay of 0.003. Regard- ing the hyperparameter configuration, the model was trained for a maximum of 100 epochs with a batch size of 8. We employed the Adam optimizer with an initial learning rate of 1× 10 −3 , coupled with a ReduceLROnPlateau scheduler. Fur- thermore, we implemented label smoothing of 0.2 to improve the model’s calibration and robustness. Through extensive ex- perimentation, we determined the optimal hyperparameters for graph construction: a sliding window size of n = 3 and an edge addition threshold of τ c = 0.3 for the co-occurrence graph, and a threshold of τ s = 0.8 for the semantic graph, which collec- tively yielded the best performance in our final results. 3.2. Main results and discussions As illustrated in Table 1, our proposed framework demonstrates superior performance, consistently outperforming all baseline methods across all metrics. Specifically, it achieves a 5-fold cross-validation accuracy of 88.81% on the training set and an accuracy of 90.00% on the test set. Table 1: Performance comparison of different methods on 5-fold cross-validation and test set (%). Method 5-Fold Cross-Validation ResultsTest Set Results AccuracyF1-ScoreRecallPrecisionAccuracyF1-ScoreRecallPrecision Luz et al. [28]78.2277.9673.4084.1779.7178.1273.5383.33 Balagopalan et al. [11]80.7882.2483.7981.2882.8683.3385.7181.08 Ajroudi et al. [12]85.0885.5883.8688.4381.4381.1680.0082.35 Cai et al. [14]84.3885.5486.1485.1084.2983.5880.0087.50 Ortiz-Perez et al. [29]86.8886.7187.3186.7981.4380.6077.1484.38 Ours88.8188.8188.8189.4390.0089.8688.5791.18 Compared to the baseline established by Luz et al. , our model demonstrates a substantial performance leap. In con- trast to Ajroudi et al., who utilized OpenAI’s text embeddings as high-dimensional features for direct classification, our approach employs BERT to establish a rigorous word-to-node mapping via token averaging. This strategy leverages deep semantic rep- resentations while simultaneously preserving the discrete sym- bolic structure of the text. Unlike the approach by Balagopalan et al., which utilizes BERT hidden features for direct classifica- tion, our multi-view gated graph attention network framework enables structural linguistic modeling by characterizing speech patterns through complementary perspectives. While Cai et al. focused on dependency or dynamic graphs, our integration of a PMI-based co-occurrence graph effectively captures temporal anomalies in speech flow. Pathological mark- ers common in AD patients—such as disjointed logical tran- sitions or labored word retrieval—are explicitly manifested through abnormal fluctuations in PMI weights. Simultaneously, the dependency graph models the integrity of the linguistic ”skeleton,” while the semantic graph utilizes cosine similarity to detect subtle semantic drift. The method proposed by Ortiz-Perez et al. achieved state- of-the-art performance under the cross-validation setting. How- ever, it was observed that this configuration leads to a signifi- cant performance degradation when the model is evaluated on the test set. This phenomenon may be attributed to the model overfitting to its respective validation split in each fold, yielding inflated overall metrics. For a more stringent assessment of per- formance, we evaluated the proposed approach against the test set, yielding an accuracy rate of 81.43%. 3.3. Ablation Study To evaluate the individual contribution of each architectural component to the diagnostic performance of our proposed method, we conducted a series of ablation experiments on the test set. By systematically removing specific graph views and the fusion mechanism, we generated five ablation variants for comparison. The results are summarized in Table 2. The most significant performance degradation occurred in the w/o Graph Structure variant, which relies solely on the aver- age pooling of BERT embeddings without any topological mod- eling. This variant achieved the lowest scores across all metrics (Acc: 81.43%, F1: 80.60%), underscoring the critical impor- tance of capturing non-linear structural dependencies in patho- logical speech. Among the specific linguistic views, the removal of the Co- occurrence Graph led to a substantial decline in the F1-score to 85.29%. This confirms that the temporal ”flow” and logi- cal progression captured via PMI are indispensable biomarkers Table 2: Ablation study results on test set (%). Model VariantAccuracyF1-Score w/o Semantic Graph85.7186.11 w/o Dependency Graph87.1486.57 w/o Co-occurrence Graph85.7185.29 w/o Gated Fusion87.1487.32 w/o Graph Structure81.4380.60 Ours90.0089.86 for Alzheimer’s Disease. Similarly, the removal of the Seman- tic Graph and Dependency Graph resulted in accuracy drops to 85.71% and 87.14%, respectively. These results underscore that while semantic content provides the ”substance” of the diagno- sis, the syntactic ”skeleton” provides the necessary linguistic constraints for robust classification. The w/o Gated Fusion variant, which utilizes a static aver- aging of the three graph outputs instead of the dynamic weight- ing network, saw its accuracy fall to 87.14%. This confirms that a ”one-size-fits-all” approach to linguistic feature integration is insufficient for AD detection. The gated network’s ability to adaptively prioritize different linguistic views based on individ- ual patient speech patterns—such as prioritizing syntactic errors in one patient and semantic incoherence in another—is a critical factor in achieving the full model’s accuracy of 90.00%. 4. Conclusion Experimental evaluations on the ADReSSo 2021 dataset vali- date that the Multi-View Gated Graph Attention Network pro- vides a robust framework for identifying AD by holistically characterizing the ”content-structure-flow” features of sponta- neous speech. One of the primary contributions of this study is the pivotal role played by the co-occurrence graph; by uti- lizing normative PMI to quantify discourse logic, it success- fully captures the incoherent logical transitions and repetitive patterns characteristic of cognitive decline that traditional se- quential models are typically unable to detect. Ablation results indicate that removing this ”flow” dimension significantly im- pairs diagnostic performance, confirming that narrative progres- sion is a crucial feature. Furthermore, unlike conventional static fusion strategies, our gated network adaptively assigns weights to linguistic features on a per-sample basis, effectively mitigat- ing the challenges posed by the inherent clinical heterogeneity of AD. Future research will prioritize evolving the framework into a comprehensive multimodal system by integrating acous- tic prosody as additional graph-based perspectives. 5. Acknowledgments This work was supported by the National Natural Science Foun- dation of China under Grant U23B2053 and by the National Talent Program under Grants E4G008, E55304, E43301, and E476. 6. Generative AI Use Disclosure During the preparation of this manuscript, we used generative AI tools to polish the English language and improve readabil- ity. These tools were not used to generate any scientific claims, experimental results, or significant parts of the manuscript. 7. References [1] H. Lindsay, J. Tr ̈ oger, and A. K ̈ onig, “Language Impairment in Alzheimer’s Disease—Robust and Explainable Evidence for AD- Related Deterioration of Spontaneous Speech Through Multilin- gual Machine Learning,” Frontiers in Aging Neuroscience, vol. Volume 13 - 2021, 2021. [2] M. Gumus, M. Koo, C. M. Studzinski, A. Bhan, J. Robin, and S. E. Black, “Linguistic changes in neurodegenerative diseases relate to clinical symptoms,” Frontiers in Neurology, vol. 15, p. 1373341, Mar. 2024. [3] X. Qi, Q. Zhou, J. Dong, and W. Bao, “Noninvasive automatic detection of Alzheimer’s disease from spontaneous speech: a re- view,” Frontiers in Aging Neuroscience, vol. 15, p. 1224723, Aug. 2023. [4] Y. Qiao, X. Yin, D. Wiechmann, and E. Kerz, “Alzheimer’s Dis- ease Detection from Spontaneous Speech Through Combining Linguistic Complexity and (Dis)Fluency Features with Pretrained Language Models,” in Interspeech 2021.ISCA, Aug. 2021, p. 3805–3809. [5] M. Kurdi, “Automatic diagnosis of alzheimer’s disease using lexi- cal features extracted from language samples,” Journal of Medical Artificial Intelligence, vol. 7, p. 13–13, Jun. 2024. [6] Z. S. Syed, M. S. S. Syed, M. Lech, and E. Pirogova, “Tack- ling the ADRESSO Challenge 2021: The MUET-RMIT System for Alzheimer’s Dementia Recognition from Spontaneous Speech,” in Interspeech 2021. ISCA, Aug. 2021, p. 3815–3819. [7] R. Shankar, Z. Goh, F. Devi, and Q. Xu, “A systematic review of explainable artificial intelligence methods for speech-based cog- nitive decline detection,” npj Digital Medicine, vol. 8, no. 1, p. 724, Nov. 2025. [8] L. Calz ` a, G. Gagliardi, R. Rossini Favretti, and F. Tamburini, “Linguistic features and automatic classifiers for identifying mild cognitive impairment and dementia,” Computer Speech & Lan- guage, vol. 65, p. 101113, Jan. 2021. [9] A. Balagopalan and J. Novikova, “Comparing Acoustic-Based Approaches for Alzheimer’s Disease Detection,” in Interspeech 2021. ISCA, Aug. 2021, p. 3800–3804. [10] X. Wei, B. Wen, Y. Lin, K. Li, M. gu, X. Wang, L. Wang, and J. Dang, “Breaking Data Efficiency Dilemma: A Federated and Augmented Learning Framework For Alzheimer’s Disease Detec- tion via Speech,” Feb. 2026, arXiv:2602.14655 [cs]. [11] A. Balagopalan, B. Eyre, F. Rudzicz, and J. Novikova, “To BERT or not to BERT: Comparing Speech and Language- Based Approaches for Alzheimer’s Disease Detection,” in Interspeech 2020.ISCA, Oct. 2020, p. 2167–2171. [On- line]. Available: https://w.isca-archive.org/interspeech 2020/ balagopalan20interspeech.html [12] K. Ajroudi, M. I. Khedher, O. Jemai, and M. A. EI-Yacoubi, “Ex- ploring the Efficacy of Text Embeddings in Early Dementia Di- agnosis from Speech,” in 2024 16th International Conference on Human System Interaction (HSI), Jul. 2024, p. 1–6, iSSN: 2158- 2254. [13] K. Chlasta, P. Struzik, and G. M. W ́ ojcik, “Enhancing dementia and cognitive decline detection with large language models and speech representation learning,” Frontiers in Neuroinformatics, vol. 19, Dec. 2025. [14] H. Cai, X. Huang, Z. Liu, W. Liao, H. Dai, Z. Wu, D. Zhu, H. Ren, Q. Li, T. Liu, and X. Li, “Exploring Multimodal Approaches for Alzheimer’s Disease Detection Using Patient Speech Transcript and Audio Data,” Jul. 2023, arXiv:2307.02514 [eess]. [15] A. E. Hallani, A. Chakhtouna, and A. Adib, “Graph Attention Networks with Dual-Edge Connectivity for Alzheimer’s Disease Detection from Speech,” IEEE Transactions on Artificial Intelli- gence, p. 1–11, 2025. [16] E. Burke, J. Gunstad, and P. Hamrick, “Comparing global and local semantic coherence of spontaneous speech in persons with Alzheimer’s disease and healthy controls,” Applied Corpus Lin- guistics, vol. 3, no. 3, p. 100064, Dec. 2023. [17] K. C. Fraser, J. A. Meltzer, and F. Rudzicz, “Linguistic Features Identify Alzheimer’s Disease in Narrative Speech,” Journal of Alzheimer’s Disease, vol. 49, no. 2, p. 407–422, Nov. 2015. [18] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. Mcleavey, and I. Sutskever, “Robust speech recognition via large-scale weak su- pervision,” in Proceedings of the 40th International Conference on Machine Learning, ser. Proceedings of Machine Learning Re- search, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, Eds., vol. 202.PMLR, 23–29 Jul 2023, p. 28 492–28 518. [19] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre- training of Deep Bidirectional Transformers for Language Under- standing,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguis- tics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio, Eds.Minneapo- lis, Minnesota: Association for Computational Linguistics, Jun. 2019, p. 4171–4186. [20] Y. Chen, L. Wu, and M. J. Zaki, “Iterative Deep Graph Learn- ing for Graph Neural Networks: Better and Robust Node Embed- dings,” Oct. 2020, arXiv:2006.13009 [cs]. [21] M. Honnibal and I. Montani, spaCy 2: Natural language un- derstanding with Bloom embeddings, convolutional neural net- works and incremental parsing, 2017, software available from https://spacy.io/. [22] O. Ivanova, I. Mart ́ ınez-Nicol ́ as, E. Garc ́ ıa-Pi ̃ nuela, and J. J. G. Meil ́ an, “Defying syntactic preservation in Alzheimer’s disease: what type of impairment predicts syntactic change in dementia (if it does) and why?” Frontiers in Language Sciences, vol. 2, p. 1199107, Aug. 2023. [23] Z. Lian and Z. Wang, “Dependency Grammar Approach to the Syntactic Complexity in the Discourse of Alzheimer Patients,” Behavioral Sciences, vol. 15, no. 10, Sep. 2025. [24] K. W. Church and P. Hanks, “Word Association Norms, Mu- tual Information, and Lexicography,” Computational Linguistics, vol. 16, no. 1, p. 22–29, 1990. [25] L. Yao, C. Mao, and Y. Luo, “Graph Convolutional Networks for Text Classification,” Nov. 2018, arXiv:1809.05679 [cs]. [26] P. Veli ˇ ckovi ́ c, G. Cucurull, A. Casanova, A. Romero, P. Li ` o, and Y. Bengio, “Graph Attention Networks,” Feb. 2018, arXiv:1710.10903 [stat]. [27] J. Arevalo, T. Solorio, M. Montes-y G ́ omez, and F. A. Gonz ́ alez, “Gated Multimodal Units for Information Fusion,” Feb. 2017, arXiv:1702.01992 [stat]. [28] S. Luz, F. Haider, S. D. L. Fuente, D. Fromm, and B. MacWhin- ney, “Detecting Cognitive Decline Using Speech Only: The ADReSSo Challenge,” in Interspeech 2021.ISCA, Aug. 2021, p. 3780–3784. [29] D. Ortiz-Perez, M. Benavent-Lledo, J. Rodriguez-Juan, J. Garcia- Rodriguez, and D. Tom ́ as, “CogniAlign: Word-level multimodal speech alignment with gated cross-attention for Alzheimer’s de- tection,” Knowledge-Based Systems, vol. 329, p. 114264, Nov. 2025.