Paper deep dive
HASS: Hierarchical Simulation of Logopenic Aphasic Speech for Scalable PPA Detection
Harrison Li, Kevin Wang, Cheol Jun Cho, Jiachen Lian, Rabab Rangwala, Chenxu Guo, Emma Yang, Lynn Kurteff, Zoe Ezzes, Willa Keegan-Rodewald, Jet Vonk, Siddarth Ramkrishnan, Giada Antonicelli, Zachary Miller, Marilu Gorno Tempini, Gopala Anumanchipalli
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/31/2026, 1:31:15 AM
Summary
HASS (Hierarchical Aphasic Speech Simulation) is a novel, clinician-guided framework designed to address data scarcity in primary progressive aphasia (PPA) research. By modeling the logopenic variant (lvPPA) through a two-layer production system—lexical retrieval impairment and phonological encoding disruption—HASS generates severity-controlled synthetic speech. Experiments demonstrate that models trained on HASS-augmented data outperform those trained on limited clinical datasets, showing superior robustness and cross-site generalization.
Entities (5)
Relation Signals (3)
HASS → simulates → lvPPA
confidence 100% · HASS aims to simulate behaviors of logopenic variant of PPA (lvPPA)
HASS → improves → PPA classification
confidence 95% · We demonstrate that HASS-generated data improves automated PPA classification
HASS → uses → VITS
confidence 95% · We synthesize speech with TTS (VITS) while explicitly preserving dysfluency.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Building a diagnosis model for primary progressive aphasia (PPA) has been challenging due to the data scarcity. Collecting clinical data at scale is limited by the high vulnerability of clinical population and the high cost of expert labeling. To circumvent this, previous studies simulate dysfluent speech to generate training data. However, those approaches are not comprehensive enough to simulate PPA as holistic, multi-level phenotypes, instead relying on isolated dysfluencies. To address this, we propose a novel, clinically grounded simulation framework, Hierarchical Aphasic Speech Simulation (HASS). HASS aims to simulate behaviors of logopenic variant of PPA (lvPPA) with varying degrees of severity. To this end, semantic, phonological, and temporal deficits of lvPPA are systematically identified by clinical experts, and simulated. We demonstrate that our framework enables more accurate and generalizable detection models.
Tags
Links
- Source: https://arxiv.org/abs/2603.26795v1
- Canonical: https://arxiv.org/abs/2603.26795v1
Trouble viewing inline? Open PDF directly →
Full Text
28,419 characters extracted from source content.
Expand or collapse full text
HASS: Hierarchical Simulation of Logopenic Aphasic Speech for Scalable PPA Detection Harrison Li 1 , Kevin Wang 1 , Cheol Jun Cho 1 , Jiachen Lian 1 , Rabab Rangwala 2 , Chenxu Guo 3 , Emma Yang 4 , Lynn Kurteff 2 , Zoe Ezzes 2 , Willa Keegan-Rodewald 2 , Jet Vonk 2 , Siddarth Ramkrishnan 2 , Giada Antonicelli 5 , Zachary Miller 2 , Marilu Gorno Tempini 2 , Gopala Anumanchipalli 1 1 UC Berkeley, USA 2 UCSF, USA 3 Zhejiang University, China 4 Columbia University, USA 5 Basque Center on Cognition, Brain and Language, Spain liharrison@berkeley.edu, kwang3170@berkeley.edu, gopala@berkeley.edu Abstract Building a diagnosis model for primary progressive aphasia (PPA) has been challenging due to the data scarcity. Collect- ing clinical data at scale is limited by the high vulnerability of clinical population and the high cost of expert labeling. To cir- cumvent this, previous studies simulate dysfluent speech to gen- erate training data. However, those approaches are not compre- hensive enough to simulate PPA as holistic, multi-level pheno- types, instead relying on isolated dysfluencies. To address this, we propose a novel, clinically grounded simulation framework, Hierarchical Aphasic Speech Simulation (HASS). HASS aims to simulate behaviors of logopenic variant of PPA (lvPPA) with varying degrees of severity. To this end, semantic, phonolog- ical, and temporal deficits of lvPPA are systematically identi- fied by clinical experts, and simulated. We demonstrate that our framework enables more accurate and generalizable detection models. Code: https://github.com/haribary/HA S Index Terms:Primary progressive aphasia, pathological speech simulation, dysfluency modeling 1. Introduction Primary progressive aphasia (PPA) is a neurodegenerative dis- order characterized by progressive language impairment. Be- cause the hallmark of PPA is progressive deterioration of lan- guage function, connected (spontaneous) speech provides rich diagnostic information for PPA variant characterization and is widely used in clinical and research assessment [1, 2, 3]. As speech-based machine learning methods advance, there is growing interest in automated screening frameworks for PPA [4, 5, 6]. However, such approaches are constrained by the lim- ited availability of high-quality naturalistic speech datasets from clinically characterized PPA patients. Data collection typically requires expert diagnosis, structured elicitation protocols, and careful annotation under strict ethical and privacy constraints [7, 8, 9]. Public resources such as DementiaBank and Aphasi- aBank provide invaluable connected-speech samples, but re- main limited in size and institutional representation, constrain- ing model development and cross-corpus robustness [7, 8, 9]. Synthetic data generation offers a potential solution, but prior dysfluency simulation efforts have largely focused on in- jecting isolated dysfluency events (e.g., repetitions, insertions, or pauses) into otherwise fluent speech [10, 11, 12, 13, 14, 15]. However, behavioral deficits in some PPA variants arise from disruptions on multiple levels (phonological, word and content) of speech [3]. Phonological deficits simulated in isolation fail to capture these interactions. As a result, such simulations of- ten lack clinical plausibility and fail to model simultaneous dis- ruptions across phoneme, lexical, and content levels. Further- more, most simulations using LLM-based or agentic generation pipelines are not grounded in the production mechanisms of a specific clinical phenotypes, limiting their utility because mod- els trained on such data may learn generic surface-level dys- fluency patterns rather than disorder-specific impairment signa- tures [12, 15, 16, 17]. To date, no end-to-end simulation frame- work has explicitly modeled a neurodegenerative language dis- order using clinically grounded production mechanisms. Among PPA variants, lvPPA is characterized by impaired word retrieval that cascades into phonological errors and halting speech [3, 18], making it an ideal test case for a multi-level sim- ulation framework. Thus, we introduce the Hierarchical Apha- sic Speech Simulation (HASS), a clinically grounded hierarchi- cal simulation framework for logopenic variant PPA (lvPPA) that models clinically defined impairment mechanisms in a two- layer production model: (1) a lexical retrieval impairment layer that generates severity-conditioned content-level disruptions and (2) a phonological encoding disruption layer that introduces severity-conditioned phoneme-level errors on a word-aligned representation. HASS produces severity-controlled synthetic lvPPA speech with co-occurring content-level and phonological dysfluencies, alongside matched controls generated using the same pipeline but without impairment injection. This enables scalable augmentation of low-resource clinical datasets while ensuring classifier differences are attributable to the simulated disorder rather than synthesis artifacts. Our experiments show that HASS improves the performance of the diagnosis model during evaluation. We will release all data, models, and code to support reproducibility. We summarize our contributions as follows: • We propose HASS, an exclusive clinician-guided simulation pipeline and the first framework to model a neurodegener- ative aphasia (lvPPA) as a holistic, structural disease rather than a collection of isolated, disconnected dysfluencies. • We introduce a scalable recipe for clinical data augmenta- tion, releasing a comprehensive, severity-controlled synthetic dataset that accurately reflects the multi-level impairments of PPA. arXiv:2603.26795v1 [eess.AS] 25 Mar 2026 • We demonstrate that HASS-generated data improves auto- mated PPA classification, with classifiers trained on HASS speech outperforming those trained on clinical recordings. We evaluate both in-domain capability and cross-site gener- alization. 2. Simulation 2.1. Dysfluent Text Generation We introduce a two-layer dysfluent text generation pipeline de- signed to model language production deficits characteristic of Logopenic Variant primary progressive aphasia (lvPPA). In par- ticular, we recruit an LLM (Gemini 3 [19]) to simulate patho- logical behaviors with detailed, clinically guided instructions. The simulator encodes clinically defined lvPPA symptoms, and was developed with SLP oversight. LvPPA’s primary impair- ment is in lexical retrieval and its downstream phonological consequences, which inspired the factorization of dysfluency modeling into content-level and phoneme-level processes. During generation, the LLM is constrained by several clin- ically grounded rules: • Lexical Bias: In lvPPA, content and phonological dysfluen- cies are correlated. Dysfluencies are applied non-uniformly, heavily biasing toward high lexical-demand loci such as low- frequency content words and multisyllabic targets. • Syntactic Adherence: Disruptions are constrained to plau- sible syntactic and discourse boundaries (e.g., clause bound- aries or pre-content-word positions). • Phenotype Exclusion:Features characteristic of non- fluent/agrammatic and semantic PPA variants (e.g., persistent agrammatism, apraxia-of-speech-like distortions, or semanti- cally empty fluent speech) are explicitly penalized to prevent phenotype drift. 2.1.1. Word Level Non-pathological spontaneous text is first generated using a di- verse set of prompts modelled in the style of connected speech questions sourced from the Quick Aphasia Battery[1]. The out- put is directed into two pipelines: the synthetic control pipeline data are injected with naturalistic dysfluencies on the word level and the synthetic dysfluency pipeline where this text now serves as the ground truth. In the dysfluency pipeline, ground-truth text is first passed to the word-level (content) dysfluency layer. Conditioned on the chosen severity variable, we instruct an LLM to introduce lvPPA-like lexical retrieval phenomena, including circumlocu- tions, false starts, and filled pauses, while preserving the in- tended message. The resulting word-level dysfluent text is then converted to a word-aligned IPA representation using [20]. Stress markers and word boundaries are retained to enable consistent alignment for downstream phonological editing and speech synthesis. 2.1.2. Phoneme Level The phoneme-level layer edits the word-aligned IPA sequence, conditioned on both the word-level dysfluent text and its IPA target form. We instruct LLMs to insert inline markers for six error types organized in a clinically motivated hierarchy. The three primary markers are [PAU] (pause insertion), [SUB] (phoneme substitution), and [DEL] (phoneme deletion), re- flecting the dominant temporal and phonological disruptions in lvPPA, where slow speaking rate, word-finding halts, and phonological paraphasias predominate [21, 18, 3, 22, 23]. Next, two secondary markers are also modelled: [REP] (sound/syllable repetition), modeled as a byproduct of self- repair during failed retrieval [3], [PRO] (phoneme prolonga- tion), reflecting mild hesitation-related lengthening [3]. Fi- nally, [INS] (phoneme insertion), is treated as a rare dysflu- ency since it is reported at lower rates than other phonologi- cal paraphasias [24]. Furthermore, marker rates are severity- conditioned and biased toward content words (≥80%), with disruption probability increasing with word length and sylla- ble complexity. Repetition is restricted to repair contexts, and within-word marker co-occurrence is capped at higher severi- ties to maintain a realistic density of dysfluencies. The output is a marked IPA string used for subsequent speech synthesis. 2.2. Synthesis of Dysfluent Speech We synthesize speech with TTS (VITS)[25] while explicitly preserving dysfluency. Phoneme-level markers ([DEL], [SUB], [INS], [REP]) are applied upstream during IPA generation, while [PAU] is realized by inserting a silence segment during and [PRO] by prolonging the target phoneme during inference. We provide both sentence-level audio outputs and concatenated utterances; sentence audio is concatenated downstream using a 50 ms crossfade. 3. Data 3.1. Marker Distribution Analysis Figure 2 confirms that the generated data respect the clinical marker hierarchy. Across all severity levels, pause, deletion, and substitution account for the majority of dysfluency events, while prolongation and repetition occur less frequently and in- sertion remains rare. Distributions shift toward higher counts with increasing severity: summing mean counts across mark- ers yields T mild = 10.0, T mod = 21.1, and T sev = 29.0 markers per file (2.1× and 2.9× increases relative to mild). The three primary markers account for 75.0% of mild events (7.5/10.0), 64.0% of moderate (13.5/21.1), and 65.5% of se- vere (19.0/29.0). At the extremes, insertion averages fewer than one event per file even at severe (μ INS = 0.0, 0.4, 0.8), whereas the dominant markers reach substantially higher rates (μ DEL = 7.1, μ SUB = 5.8, μ PAU = 6.1 at severe). 3.2. Dataset Composition The HASS corpus comprises 4,773 sentence-level clips to- talling 12.81 hours of synthesized audio. Of these, 2,007 are control utterances and 2,766 are dysfluent (871 mild, 1,101 moderate, 794 severe). Speech is synthesized using 95 speakers from the VCTK corpus (out of 109 available voices) across 40 unique ground-truth prompts. Controls are generated through the same synthesis pipeline, speakers, and prompts but without lvPPA-specific dysfluency injection, ensuring that any classifier differences are attributable to the simulated impairment rather than speaker or synthesis artifacts. 4. Experiments and Results To validate the clinical utility of our synthetic corpus, we eval- uate a classification model trained on HASS-generated data and assess its zero-shot generalization on real-world clinical record- ings. We evaluate on real lvPPA patient audio from the Baycrest Figure 1: Overview of the HASS hierarchical simulation pipeline. Table 1: Example output across severity levels for a single ground-truth sentence. Text: word-level dysfluent output. IPA: phoneme- level output with inline markers. Ground truth: “The house would go completely dark, save for the single amber glow of a hearth fire.” SeverityOutput CONTROLText: The house would go completely dark, save for the single amber glow of a hearth fire. IPA: D@ h"aUs wUd g"oU k@mpl"i`utli d"A`uˆok, s"eIv f3ˆoD@ s"INg@l "æmb3ˆo gl"oU @v@ h"A`uˆoh f"aI@ˆo MILDText: The house would go comple[DEL]ly dark, except for the single, you know, the ori[SUB]mge light of the, the place where you burn the wood, the hearth fire. IPA: D@ h"aUs wUd g`oU k@mpl"i`u[DEL]li d"A`uˆok, Eks"Ept f3ˆoD@ s"INg@l, ju`u n"oU, DI "Oˆoim[SUB]dZ l"aIt VvD@, D@ pl"eIs w`Eˆo ju`u b"3`un D@ w"Ud, D@ h"A`uˆoh f"aI@ˆo MODERATEText: The house would go, it would go comple[DEL]ly dark, except for the, the oran[SUB]z light, the am[DEL]ber glow from the, the place where you burn the wood, [PAU] the hea[DEL]th. IPA: D@ h"aUs wUd g"oU, It wUd g`oU k@mpl"i`u[DEL]li d"A`uˆok, Eks"Ept f3ˆoD@, DI "OˆoIndz[SUB] l"aIt, DI "æm[DEL]b3ˆo gl"oU fˆo VmD@, D@ pl"eIs w`Eˆo ju`u b"3`un D@ w"Ud, [PAU] D@ h"A`u[DEL]T SEVEREText: It was, it went, uh, [PAU] no ligh[DEL], just the, the one, the fi[PRO]re thin[SUB], in. . . inside[REP] IPA: It w"Vz, It w"Ent, "V, [PAU] n"oU l"aI[DEL], dZ"Vst D@, D@ w"Vn, D@ f"aI[PRO]@ˆo T"In[SUB], I. . . Ins"aId [REP] Table 2: Cross-site performance comparison (mean± std). ModelAUCF1Recall (Dys) Baseline0.850± 0.1220.778± 0.1650.659± 0.238 HASS0.892± 0.0760.800± 0.0720.899± 0.066 PPA Protocol corpus [26] and the Hopkins PPA corpus [27] in DementiaBank, and on two control datasets: the Delaware cor- pus [9] from DementiaBank and the Capilouto corpus [28] from AphasiaBank. Both control datasets are selected using explicit exclusion criteria (e.g., no neurological or cognitively deterio- rating conditions, fluent English, and no clinically significant depression). Baycrest, Delaware, and Capilouto share the stan- dard TalkBank discourse protocol [7, 8, 9], whereas Hopkins uses a clinical assessment battery [27] comprising naming, pas- sage reading, counting, and story retelling. Classifiers are eval- uated in a strictly cross-site design: trained on Baycrest lvPPA and Delaware controls and tested on Hopkins dysfluent and Capilouto controls, then vice versa. This protocol and recording mismatch across four independent sites provides a stringent test of cross-corpus generalization. 4.1. Modeling details We fine-tune Wav2Vec 2.0 [29] with the base model size using Low-Rank Adaptation (LoRA) [30]. We apply LoRA adapters exclusively to the query and value projection layers (Q proj , V proj ). We compare two distinct training regimes: • Baseline: Due to the limited size of the real clinical datasets, we employ 5-fold cross-validation to ensure reliable per- formance estimates. Data is grouped by speaker to pre- vent train/test contamination. To mitigate confounding back- ground variability, all baseline train data is enhanced with MossFormer2 SE48K [31] prior to feature extraction. • HASS: This model is trained using all samples from our gen- erated HASS corpus. We sample fixed-length 15 s windows (240,000 samples at 16 kHz) from concatenated synthetic au- dio, which is grouped by ground-truth prompt, severity, and speaker. We apply on-the-fly augmentation with speed perturbation (0.9– 1.1), additive Gaussian noise (10–20 dB SNR), volume jit- ter (±6 dB), and reverberation via synthetic room impulse re- sponses (RT60 0.2–0.8 s), with per-sample probabilities of 0.5, 0.5, 0.5, and 0.3, respectively. Each test fold is balanced to the minority class and evaluated on pre-enhanced audio. We report AUC-ROC, macro F1, and recall metrics all five folds. 4.2. Cross-Site Model Evaluation To further evaluate the advantage of our simulated data, we compare models trained on synthetic speech against baseline models in a strict cross-site scenario. In particular, we evalu- ate generalization by training a model on data from one clinical site and testing it on another. We partition our real datasets into Figure 2: Distribution of phonological dysfluency markers across severity levels.Dashed lines indicate per-severity means. two evaluation domains: Domain A comprises Baycrest (dys- fluent) and Delaware (control), while Domain B comprises JHU (dysfluent) and Capilouto (control). We separate Baycrest and JHU into different domains due to their contrasting protocols for cross-domain evaluation. All models are trained with the same setting as §4.1 using Wav2Vec 2.0 base. • Baseline: Trained on real patient recordings and controls. We train a separate baseline model for each cross-site scenario (i.e., trained on Domain A and tested on Domain B, and vice versa). • HASS: A single model is trained exclusively on HASS- generated streams. This model is then evaluated on the real cross-site test domains. 4.3. Results HASS provides synthetic samples that enable more accu- rate automatic lvPPA diagnosis Table 2 summarizes perfor- mance of the comparison of HASS-trained model and the base- line models The HASS model outperforms the baseline mod- els which only use limited real-world data across all primary metrics. Figure 3 illustrates this advantage, as a HASS-trained model not only achieves higher mean AUC-ROC, but displays tighter variance across folds, indicating a more stable learning signal. HASS-trained models demonstrate robust cross-site generalization. We evaluated cross-site classification capabil- ity of HASS-trained model. As illustrated in Figure 4, the model trained exclusively on HASS-generated data outperforms base- line models trained on real clinical recordings when evaluated on recordings from different clinical sites. Notably, the HASS- trained model achieves a higher AUC-ROC in a strict cross-site setting (trained on one institutional dataset and tested on an- other, and vice versa). Cross-site robustness is essential for real- world clinical utility. Overfitting to local variations in record- ing environments and elicitation protocols remains the primary Figure 3:Comparison of ROC curves for LoRA models trained on real vs HASS speech, evaluated using 5-fold cross- validation. Mean ROC curve and ±1 standard deviation shad- ing. Figure 4: Cross-site ROC curves for LoRA SFT on w2v2-base. Test data as follows - Group 1: JHU + Capilouto. Group 2: Baycrest + Delaware. barrier to deploying universal, scalable diagnostic models. This highlights two key advantages of our synthetic approach: gener- alization and scalability. HASS provides a mechanism to gener- ate virtually limitless, highly diverse training samples that cap- ture the core clinical phenotype of lvPPA without overfitting to the acoustic artifacts or demographic biases of a single clinical site. 5. Conclusion In this work, we introduced HASS, a novel, clinician-guided hi- erarchical simulation framework designed to address the critical data scarcity bottleneck in automated primary progressive apha- sia (PPA) screening. By explicitly modeling the complex, multi- level language deficits characteristic of the logopenic variant (lvPPA) at both the lexical and phonological levels, HASS gen- erates highly realistic, severity-controlled synthetic speech. Our empirical evaluations demonstrate that the diagnostic model trained on HASS-generated data not only outperform base- lines trained strictly on real clinical recordings, but also exhibit superior robustness and generalization in stringent, cross-site evaluations. Ultimately, this framework provides a scalable, privacy-preserving pathway for augmenting low-resource clini- cal datasets, paving the way for more reliable and robust speech- based diagnostic tools for neurodegenerative diseases. Limitations. Standard phoneme-to-speech architectures like VITS are inherently optimized for fluent speech and may not perform when forced to generate severe phonological errors. Furthermore, the neurological variability of dysfluent speech remains highly complex and under active study; because indi- vidual symptoms progress heterogeneously, true PPA severity manifests in more nuanced ways. 6. References [1] S. M. Wilson, D. K. Eriksson, S. M. Schneck, and J. M. Lucanie, “A quick aphasia battery for efficient, reliable, and multidimen- sional assessment of language function,” PLoS One, vol. 13, no. 2, p. e0192773, Feb. 2018. [2] J. A. Matias-Guiu, P. Su ́ arez-Coalla, M. Yus, V. Pytel, L. Hern ́ andez-Lorenzo, C. Delgado-Alonso, A. Delgado- ́ Alvarez, N. G ́ omez-Ruiz, C. Polidura, M. N. Cabrera-Mart ́ ın, J. Mat ́ ıas- Guiu, and F. Cuetos, “Identification of the main components of spontaneous speech in primary progressive aphasia and their neu- ral underpinnings using multimodal MRI and FDG-PET imag- ing,” Cortex, vol. 146, p. 141–160, Jan. 2022. [3] M. L. Gorno-Tempini, A. E. Hillis, S. Weintraub, A. Kertesz, M. Mendez, S. F. Cappa, J. M. Ogar, J. D. Rohrer, S. Black, B. F. Boeve, F. Manes, N. F. Dronkers, R. Vandenberghe, K. Rascov- sky, K. Patterson, B. L. Miller, D. S. Knopman, J. R. Hodges, M. M. Mesulam, and M. Grossman, “Classification of primary progressive aphasia and its variants,” Neurology, vol. 76, no. 11, p. 1006–1014, Mar. 2011. [4] N. Rezaii, D. Hochberg, M. Quimby, B. Wong, M. Brickhouse, A. Touroutoglou, B. C. Dickerson, and P. Wolff, “Artificial intelligence classifies primary progressive aphasia from connected speech,” Brain, vol. 147, no. 9, p. 3070–3082, 06 2024. [Online]. Available: https://doi.org/10.1093/brain/awae196 [5] F. Peters, W. R. Bevan-Jones, G. Threlfall, J. M. Harris, J. S. Snowden, M. Jones, J. C. Thompson, D. J. Blackburn, and H. Christensen, “Automatic Detection and Sub-typing of Primary Progressive Aphasia from Speech: Integrating Task-Specific Fea- tures and Spatio-Semantic Graphs,” in Interspeech 2025, 2025, p. 5288–5292. [6] J. M. J. Vonk, J. Lian, Z. Ezzes, C. J. Cho, B. T. Morin, R. Bogley, Z. Miller, M. L. Mandelli, G. Anumanchipalli, and M. L. Gorno- Tempini, “Automated lexical dysfluency analysis to differentiate primary progressive aphasia variants,” in Alzheimer’s Association International Conference, 2025. [7] M. M. Forbes, D. Fromm, and B. MacWhinney, “Aphasiabank: A resource for clinicians,” Seminars in Speech and Language, 2012, pMCID: PMC4073291. [Online]. Available:https: //pmc.ncbi.nlm.nih.gov/articles/PMC4073291/ [8] B. MacWhinney, D. Fromm, M. Forbes, and A. Holland, “Aphasiabank: Methods for studying discourse,” Aphasiology, vol. 25, no. 11, p. 1286–1307, 2011, pMID: 22923879. [Online]. Available: https://doi.org/10.1080/02687038.2011.589893 [9] A. M. Lanzi, A. K. Saylor, D. Fromm, H. Liu, B. MacWhinney, and M. L. Cohen, “Dementiabank:Theoretical rationale, protocol, and illustrative analyses,” American Journal of Speech- Language Pathology, vol. 32, no. 2, p. 426–438, 2023. [Online]. Available: https://pubs.asha.org/doi/abs/10.1044/2022 AJSLP-2 2-00281 [10] J. Lian, C. Feng, N. Farooqi, S. Li, A. Kashyap, C. J. Cho, P. Wu, R. Netzorg, T. Li, and G. K. Anumanchipalli, “Unconstrained dysfluency modeling for dysfluent speech transcription and de- tection,” in 2023 IEEE Automatic Speech Recognition and Under- standing Workshop (ASRU). IEEE, 2023, p. 1–8. [11] J. Lian, X. Zhou, Z. Ezzes, J. Vonk, B. Morin, D. P. Baquirin, Z. Miller, M. L. Gorno Tempini, and G. Anumanchipalli, “Ssdm: Scalable speech dysfluency modeling,” in Advances in Neural In- formation Processing Systems, vol. 37, 2024. [12] J. Zhang, X. Zhou, J. Lian, S. Li, W. Li, Z. Ezzes, R. Bogley, L. Wauters, Z. Miller, J. Vonk, B. Morin, M. Gorno-Tempini, and G. Anumanchipalli, “Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection,” p. 1853– 1857, 2025. [13] J. Lian, X. Zhou, C. Guo, Z. Ye, Z. Ezzes, J. Vonk, B. Morin, D. Baquirin, Z. Mille, M. L. G. Tempini, and G. K. Anu- manchipalli, “Automatic detection of articulatory-based disfluen- cies in primary progressive aphasia,” 2025. [14] X. Zhou, A. Kashyap, S. Li, A. Sharma, B. Morin, D. Baquirin, J. Vonk, Z. Ezzes, Z. Miller, M. Tempini, J. Lian, and G. Anu- manchipalli, “YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection,” in Interspeech 2024, 2024, p. 937–941. [15] T. Kourkounakis, A. Hajavi, and A. Etemad, “Fluentnet: End-to- end detection of stuttered speech disfluencies with deep learning,” IEEE/ACM Transactions on Audio, Speech, and Language Pro- cessing, vol. 29, p. 2986–2999, 2021. [16] G. C. Imaezue and H. Marampelly, “Abcd: A simulation method for accelerating conversational agents with applications in aphasia therapy,” Journal of Speech, Language, and Hearing Research, vol. 68, no. 7, p. 3322–3336, 2025. [Online]. Available: https://pubs.asha.org/doi/abs/10.1044/2025 JSLHR-25-00003 [17] J. M. Pittman, A. P. Jr., Y. Medina-Santos, and B. C. Stark, “Towards a method for synthetic generation of persons with aphasia transcripts,” 2025. [Online]. Available:https: //arxiv.org/abs/2510.24817 [18] M. L. Gorno-Tempini, S. M. Brambati, V. Ginex, J. Ogar, N. F. Dronkers, A. Marcone, D. Perani, V. Garibotto, S. F. Cappa, and B. L. Miller, “The logopenic/phonological variant of primary pro- gressive aphasia,” Neurology, vol. 71, no. 16, p. 1227–1234, Oct. 2008. [19] Gemini Team Google, “Gemini: A family of highly capable mul- timodal models,” arXiv preprint arXiv:2312.11805, 2023. [20] M. Bernard and H. Titeux, “Phonemizer:Text to phones transcription for multiple languages in python,” Journal of Open Source Software, vol. 6, no. 68, p. 3958, 2021. [Online]. Available: https://doi.org/10.21105/joss.03958 [21] The Association for Frontotemporal Degeneration (AFTD), “Know the signs. . . know the symptoms: Logopenic variant ppa,” https://w.theaftd.org/wp-content/uploads/2018/03/FTD-Signs -and-Symptoms-lvPPA.pdf, 2018. [22] M. L. Henry and S. M. Grasso, “Assessment of individuals with primary progressive aphasia,” Seminars in Speech and Language, vol. 39, no. 3, p. 231–241, 2018, pMCID: PMC6464628. [Online]. Available: https://pmc.ncbi.nlm.nih.gov/articles/PMC6 464628/ [23] D. Petroi, J. R. Duffy, A. Borgert, E. A. Strand, M. M. Machulda, M. L. Senjem, C. R. Jack, K. A. Josephs, and J. L. Whitwell, “Neuroanatomical correlates of phonologic errors in logopenic progressive aphasia,” Brain and Language, vol. 204, p. 104773, 2020. [Online]. Available: https://w.sciencedirect.com/scienc e/article/pii/S0093934X20300328 [24] S. G. H. Dalton, C. Shultz, M. L. Henry, A. E. Hillis, and J. D. Richardson, “Describing phonological paraphasias in three variants of primary progressive aphasia,” Am. J. Speech. Lang. Pathol., vol. 27, no. 1S, p. 336–349, Mar. 2018. [25] J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to- speech,” in Proceedings of the 38th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139.PMLR, 18–24 Jul 2021, p. 5530–5540. [Online]. Available: https: //proceedings.mlr.press/v139/kim21f.html [26] A. Kielar, T. Deschamps, R. Jokel, and J. A. Meltzer, “Abnor- mal language-related oscillatory responses in primary progressive aphasia,” NeuroImage Clin., vol. 18, p. 560–574, Mar. 2018. [27] D. C. Tippett, C. B. Thompson, C. Demsky, R. Sebastian, A. Wright, and A. E. Hillis, “Differentiating between subtypes of primary progressive aphasia and mild cognitive impairment on a modified version of the frontal behavioral inventory,” PloS One, vol. 12, no. 8, p. e0183212, 2017. [28] G. Capilouto, “AphasiaBank english protocol capilouto corpus,” doi:10.21415/HTMN-5P65, 2026, accessed 2026-03-04. [29] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,” Advances in neural information processing systems, vol. 33, p. 12 449–12 460, 2020. [30] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al., “Lora: Low-rank adaptation of large language models.” Iclr, vol. 1, no. 2, p. 3, 2022. [31] S. Zhao, Y. Ma, C. Ni, C. Zhang, H. Wang, T. H. Nguyen, K. Zhou, J. Yip, D. Ng, and B. Ma, “Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time- domain monaural speech separation,” 2024. [Online]. Available: https://arxiv.org/abs/2312.11825