Paper deep dive
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
Tianle Yang, Chengzhe Sun, Phil Rose, Cassandra L. Jacobs, Siwei Lyu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/26/2026, 2:23:20 AM
Summary
This study introduces a segmental-level prosodic probing framework to evaluate how neural TTS models (specifically Tacotron 2 and FastSpeech 2) reproduce consonant-induced f0 perturbation. By comparing synthetic and natural speech across lexical frequency strata, the authors demonstrate that while these models can reproduce fine-grained prosodic effects for high-frequency (seen) words, they fail to generalize to low-frequency (unseen) items. This suggests that current TTS architectures rely heavily on lexical-level memorization rather than abstract segmental-prosodic encoding.
Entities (5)
Relation Signals (3)
Tacotron 2 → trainedon → LJ Speech
confidence 100% · using Tacotron 2 and FastSpeech 2 trained on the same speech corpus (LJ Speech)
FastSpeech 2 → trainedon → LJ Speech
confidence 100% · using Tacotron 2 and FastSpeech 2 trained on the same speech corpus (LJ Speech)
Neural TTS Systems → evaluatedfor → Consonant-induced f0 perturbation
confidence 95% · This study proposes a segmental-level prosodic probing framework to evaluate neural TTS models' ability to reproduce consonant-induced f0 perturbation
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This study proposes a segmental-level prosodic probing framework to evaluate neural TTS models' ability to reproduce consonant-induced f0 perturbation, a fine-grained segmental-prosodic effect that reflects local articulatory mechanisms. We compare synthetic and natural speech realizations for thousands of words, stratified by lexical frequency, using Tacotron 2 and FastSpeech 2 trained on the same speech corpus (LJ Speech). These controlled analyses are then complemented by a large-scale evaluation spanning multiple advanced TTS systems. Results show accurate reproduction for high-frequency words but poor generalization to low-frequency items, suggesting that the examined TTS architectures rely more on lexical-level memorization than on abstract segmental-prosodic encoding. This finding highlights a limitation in such TTS systems' ability to generalize prosodic detail beyond seen data. The proposed probe offers a linguistically informed diagnostic framework that may inform future TTS evaluation methods, and has implications for interpretability and authenticity assessment in synthetic speech.
Tags
Links
- Source: https://arxiv.org/abs/2603.21078v1
- Canonical: https://arxiv.org/abs/2603.21078v1
Trouble viewing inline? Open PDF directly →
Full Text
55,568 characters extracted from source content.
Expand or collapse full text
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation Tianle Yang a,∗ , Chengzhe Sun b , Phil Rose c , Cassandra L. Jacobs a , Siwei Lyu b a University at Buffalo, Department of Linguistics, Buffalo, 14260, NY, United States b University at Buffalo, Department of Computer Science and Engineering, Buffalo, 14260, NY, United States c Australian National University, Emeritus Faculty, Canberra, 0200, ACT, Australia Abstract This study proposes a segmental-level prosodic probing framework to evaluate neural TTS models’ ability to re- produce consonant-induced f0 perturbation, a fine-grained segmental-prosodic effect that reflects local articulatory mechanisms. We compare synthetic and natural speech realizations for thousands of words, stratified by lexical fre- quency, using Tacotron 2 and FastSpeech 2 trained on the same speech corpus (LJ Speech). These controlled analyses are then complemented by a large-scale evaluation spanning multiple advanced TTS systems. Results show accurate reproduction for high-frequency words but poor generalization to low-frequency items, suggesting that the examined TTS architectures rely more on lexical-level memorization rather than abstract segmental-prosodic encoding. This finding highlights a limitation in such TTS systems’ ability to generalize prosodic detail beyond seen data. The pro- posed probe offers a linguistically informed diagnostic framework that may inform future TTS evaluation methods, and has implications for interpretability and authenticity assessment in synthetic speech. Keywords: Text-to-speech synthesis, F0 perturbation, Phonetic modeling, Pitch prediction 1. Introduction In recent years, neural text-to-speech (TTS) systems have achieved remarkable naturalness, enabled by several ad- vances including end-to-end modeling (Wang et al., 2017; Kim et al., 2021), improved prosody prediction (Ren et al., 2020; Ye et al., 2023), and high-fidelity neural vocoders (Van Den Oord et al., 2016; Kong et al., 2020). While these improvements have led to more fluent and expressive synthetic speech, current evaluations have primarily focused on higher-level prosodic features. However, the ability of TTS models to replicate fine-grained phonetic effects remains underexplored. One such effect is consonant-induced f 0 perturbation (hereafter referred to as f 0 perturbation), a well- documented phenomenon in phonetics where the fundamental frequency ( f 0 ) of a vowel is systematically influenced by the immediately preceding consonant (Kirby and Ladd, 2016; Xu and Xu, 2021; Gao and Kirby, 2024; Yang and Faytak, 2025). Unlike high-level prosodic patterns, f 0 perturbation effects often arise from short-range dependencies between local segmental context, such as the voicing, aspiration, and glottal tension of a preceding consonant that are not explicitly supervised during the training of TTS models. If such perturbation effects were to emerge in synthetic speech, this would suggest that a TTS model has internalized not just surface-level statistical associations, but also the systematic co-variation between segmental features and continuous acoustic parameters. This makes perturbation an ideal linguistically grounded probe for evaluating how deeply such systems encode structured articulatory-acoustic relationships. 1.1. Background 1.1.1. Previous work on F0 perturbation To better understand what it means for a TTS system to reproduce f 0 perturbation, we briefly summarize the patterns observed in natural speech. The most widely observed consonant- induced f 0 perturbation is that voiceless ∗ Corresponding author: tianleya@buffalo.edu arXiv:2603.21078v1 [cs.CL] 22 Mar 2026 obstruents tend to induce a higher f 0 on the following vowel than their voiced counterparts. While the effect of voiceless obstruents is robustly attested across languages, the effect of voiced obstruents varies: some studies report a lowering of f 0 (Chen, 2011; Rose and Yang, 2022), others find a neutral effect (Hanson, 2009; Yang and Faytak, 2025), depending on the local and global prosody (Xu and Xu, 2003), annotation and normalization method (Xu and Xu, 2021), or language specific phonetics and phonology (Zhang, 2020). These perturbation patterns are generally attributed to well-understood physiological mechanisms. The higher f 0 after voiceless obstruents is typically attributed to increased vocal fold tension (Park and Stepp, 2019). This gesture suppresses voicing during oral closure by making the vocal folds stiffer and less able to oscillate, especially in the absence of a strong transglottal pressure gradient (Halle and Stevens, 1971; Kirby and Ladd, 2016). In contrast, the lower f 0 associated with voiced obstruents is linked to laryngeal lowering, which facilitates sustained voicing during closure (Hoole and Honda, 2011; Solé, 2018). In addition to the f 0 perturbation effect induced by the preceding consonants, the f 0 realization of a vowel is also influenced by the vowel itself. High vowels such as [i] tend to be produced with a higher f 0 compared to low vowels such as [6] (Van Hoof and Verhoeven, 2011; Jacewicz and Fox, 2015). This phenomenon, often referred to as the vowel intrinsic F0 (VF0) effect, is observed across a wide range of languages and is also thought to reflect biomechanical differences in the production of high versus low vowels (Whalen and Levitt, 1995; Whalen et al., 1995). Although the VF0 effect is not the focus of our study, it is nonetheless important to carefully model this effect so that any observed f 0 differences can be more reliably attributed to the effects of consonantal context rather than to vowel properties. In the present experiment, we controlled for vowel height by categorizing vowels according to the most recent IPA classification, grouping them into three levels: high, mid, and low. Other factors, such as the aspiration of a consonant, may also influence f 0 perturbation, but prior research suggests that these effects are relatively subtle and short-lived (Löfqvist et al., 1989; Clements and Khatiwada, 2007; Schertz and Khan, 2020) compared with the CF0 and VF0 effects described above. 1.1.2. Comparison Baseline for F0 perturbation Because the consonant-induced f 0 perturbation is by definition a relative effect, a suitable baseline is necessary to quantify and interpret any pitch differences following different consonant types. In this study, following the ap- proach originally proposed by Hanson (2009) and subsequently validated by Kirby and Ladd (2016), we adopt vowels following sonorant consonants (such as [m], [n], [N]) as our reference or baseline condition. The choice of sonorants as a reference rests primarily on physiological grounds. Sonorants such as the nasal consonant involve minimal adjustments in the supralaryngeal cavity or the state of the glottis to maintain the vibration of the vocal fold (Hanson, 2009). Thus, sonorants provide an articulatorily neutral point against which perturbations induced by other consonant types can be reliably measured. In addition, previous evidence from Kirby and Ladd (2016) further justifies the selection of sonorant-initial vowels as a baseline. They demonstrated that vowels following sonorants, especially nasals, show consistent and stable f 0 trajectories across different languages. This pattern further shows the reliability of sonorant consonants as a stable reference for quantifying onset-related perturbation. 1.2. Comparing different model structure Given the automatic nature of these acoustic results 1 , their presence in synthetic speech may depend on how sen- sitively the model encodes fine-grained segmental interaction. This consideration becomes especially relevant when comparing different TTS architectures. Autoregressive models such as Tacotron 2 (Shen et al., 2018) generate acous- tic frames sequentially, with each frame conditioned on the previous ones. This frame-level temporal recursion allows acoustic dependencies to be implicitly propagated and shaped over time. In contrast, non-autoregressive models such as FastSpeech 2 (Ren et al., 2020) generate acoustic representations in parallel and therefore lack such feedback. To compensate for this architectural constraint, temporal structure is imposed through predicted durations and positional encodings, while the f 0 contour is modeled explicitly via a dedicated pitch predictor rather than emerging implicitly 1 Note that the physiological mechanisms itself is not completely automatic, but may instead reflect some degree of speaker-level control or intentional modulation (Kingston and Diehl, 1994; Zhang, 2020). From a control-based perspective, the magnitude and duration of consonant- induced f 0 perturbation (CF0) and vowel-induced f 0 perturbation (VF0) effects may vary across languages depending on the functional need to enhance particular phonological contrasts. For instance, enhancing voicing through stronger CF0 cues, or reinforcing vowel height via more pronounced VF0 patterns. 2 through sequential decoding. These architectural differences imply distinct mechanisms by which segmental effects, such as onset-related f 0 perturbation, may be learned and generalized. To investigate whether TTS architecture influences the realization of segmental f 0 effects, we focus on two repre- sentative models: Tacotron 2 and FastSpeech 2. These two models represent distinct generation paradigms: Tacotron 2 uses an autoregressive mechanism that produces each frame conditioned on previous outputs, whereas FastSpeech 2 generates all frames in parallel. This contrast allows us to examine whether generation strategy affects the model’s ability to reproduce segmentally conditioned f 0 perturbation patterns. Another motivation for selecting these models is the availability of a consistent baseline. Both models are trained on the LJSpeech corpus (Ito and Johnson, 2017), a single-speaker dataset that provides the natural recordings (human speech) used as our ground-truth reference. This setup allows for a direct comparison between synthetic and natural realizations of the same items from the same speaker. Lastly, we also considered the accessibility of both models. Tacotron 2 and FastSpeech 2 offer open train- ing data, pretrained weights, and clear architectural documentation, which makes the results easily reproducible and interpretable. This level of accessibility contrasts with most TTS systems, which are often partially open-access or entirely proprietary. 1.3. Implication In sum, this study presents the first systematic evaluation of whether neural TTS models capture onset-induced f 0 perturbation, a segmental-level phonetic effect. We also examine whether differences in model architecture, par- ticularly the distinction between autoregressive and non-autoregressive generation, influence this ability. Our results show that neither Tacotron 2 nor FastSpeech 2 accurately captures these effects in low-frequency or unseen items. This finding is later confirmed by a large-scale study in section 5, which suggests that TTS architectures may benefit from mechanisms that more explicitly model segmental-prosodic interactions. 2. Experimental Setting We conduct two sets of experiments in this study. Experiment 1 (section 4) offers a theory-driven evaluation of whether neural TTS models (Tacotron 2 and FastSpeech 2) can reproduce segmental f 0 perturbation patterns, by directly comparing synthetic and natural speech from the same single-speaker corpus (LJ Speech), while Experiment 2 (section 5) demonstrates the practical transferability of these phonetic cues in a large-scale, multi-speaker setting, using real and fake speech to assess their robustness and diagnostic utility. 2.1. Speech generating and alignment In addition to the pre-segmented LJ Speech dataset described above, we synthesized 4,210 sentences using Fast- Speech 2 and Tacotron 2, both trained exclusively on the LJ Speech corpus. The sentences were randomly sampled from the Corpus of Contemporary American English (COCA), a balanced corpus comprising approximately one bil- lion words drawn from a range of genres, including spoken language, fiction, magazines, newspapers, and academic texts, spanning from 1990 to 2015 (Davies, 2015). The generated audio was aligned with its corresponding text using the MFA (Montreal Forced Aligner; McAuliffe et al., 2017a), which produced phoneme-level time alignments based on the pronunciation dictionary and the orthographic transcription of each audio file. To ensure the analysis was based on reliable f 0 trajectories, we excluded vowel tokens for which pitch extraction failed for more than 50% of the vowel duration (10.2%), as well as those with durations shorter than 50 ms (5.3%) according to established practice (Ting et al., 2025). Vowel segment boundaries obtained from MFA were retained as-is, unless obvious alignment errors, such as sudden changes in f 0 slope, were detected during manual inspection. For quality control, any utterances with mismatched transcripts, corrupted audio, or segmentation issues were removed prior to feature extraction. 2.2. Acoustic extraction Audio samples were aligned by MFA into a TextGrid in Praat (Boersma, 2007). For Experiment 2, which involves multiple speakers with substantially different pitch ranges, the f 0 data were extracted using the speaker-adapted al- gorithm of the Polyglot and Speech Corpus Tools (McAuliffe et al., 2017b). This algorithm involves two passes through Praat to calculate by-speaker pitch ranges for f 0 and then re-extract pitch using these ranges. To further ensure accuracy, we manually checked selected files from each speaker. 3 For Experiment 1, which compares natural speech from the LJ Speech corpus with synthetic speech generated to match the same target speaker, we intentionally did not apply speaker-specific pitch adaptation. Because both the natural and synthetic speech in this experiment originate from the same speaker, introducing different pitch parame- terizations would introduce unnecessary methodological asymmetries and reduce transparency. The analysis focused specifically on the vocalic segments of the speech signal, from which the f 0 values were time-normalized at 21 equidistant time points across the vowel duration 2 . This sampling resolution was selected to ensure sufficient temporal granularity for capturing fine-grained phonetic variation. By normalizing f 0 trajectories in time rather than using fixed frame indices, we ensure that tokens with different vowel durations can be meaningfully compared on the same relative temporal scale. To facilitate this analysis, we selected a consistent set of vowel and onset segments across systems and frequency conditions. Table 1 lists the specific vowels and obstruents included. Segment Type Selected Phones [+voi] obstruent[b], [d], [g], [dZ], [Z], [v], [D], [z] [−voi] obstruent[p], [t], [k], [tS], [f], [T], [s], [S], [h], [p h ], [t h ], [k h ], [c], [ç] sonorant[m], [n], [N], [ñ], [l], [j], [w], [m j ], [n " ] vowel[i], [i:], [I], [E], [æ], [A], [A:], [@], [6], [u], [U], [O] Table 1: IPA phones used in the analysis, grouped by segment type. This phone inventory is based on the MFA English (US) dictionary (McAuliffe and Sonderegger, 2022). 2.3. Statistical modeling We used AR1 GAMMs (generalized additive mixed models; Wood (2017); see Sóskuthy (2017) for introductions) to model the time-varying, potentially non-linear effects of onset voicing on the f 0 trajectory. A pooled model includ- ing data from all vowels following three target onset types (voiceless obstruent, voiced obstruent, sonorant) was fit using the bam() function (fREML method) in mgcv (version 1.9.1) in R (R version 4.4.0). Each model included a group-wise smooth spline term (by onset type, k=5) 3 for onset type as predictor variables, using thin plate regression splines to capture potentially non-linear effects of onset type on f 0 trajectories over time. In addition, a factor smooth (m=1, k=5) for words 4 and vowel height (m=1, k=5) were included in the overall model as control predictors, as we expect word-specific phonetics and vowel height to have effects on the f 0 realization (Gahl, 2008; Jacewicz and Fox, 2015). Difference plots for each GAMM (excluding random effects) were generated as a means of visually demonstrating where the f 0 trajectories for different onset types reached a difference significantly greater than zero (based on the default 95% confidence level). We divided our word list that contains 14,387 words in half based on their lexical frequency, and the lexical fre- quency information used to rank the word frequency was obtained from the SUBTLEX-US frequency list (Brysbaert and New, 2009). For each onset category in each speech source, we randomly selected 1,000 tokens (detailed in Table 2). The rationale for this approach is 1) we want to have balanced material (token) for each category as having bal- anced designs facilitate accurate estimation of fixed effects, reduce bias in model fitting, and ensure that smooth terms in GAMMs are estimated with adequate support. 2) to control for lexical frequency, since high-frequency words are more likely to have appeared in the training data, allowing the model to rely on surface familiarity. Low-frequency words, by contrast, test whether the model can abstract and apply prosodic patterns to unfamiliar items. 3) Incorpo- rating all the words in a corpus as a factor-smooth term is computationally intractable and unnecessary. Each level of the factor requires estimating a separate smooth, which results in high memory consumption and increased model complexity. 2 Following a widely adopted procedure in time-normalized speech analysis (e.g., Williams and Escudero, 2014; Yang and Faytak, 2025) 3 Smooth terms were initially specified with a basis dimension of k = 5, following common practice in modeling normalized f 0 trajectories in speech. This choice generally offers sufficient flexibility to capture typical contour shapes (e.g., rising, falling, or single-peaked patterns) without overfitting. We used gam.check() to assess the adequacy of smoothing. The diagnostics showed no evidence of under-smoothing (all k-index values close to 1, all p-values > 0.05), indicating that the specified basis dimensions were appropriate. 4 Smooths were parametrized with a nonlinear penalty of order 1 (m = 1) to avoid undersmoothing by speaker differences, as suggested by Wieling (2018) and Ting et al. (2025). Also note that the factor smooth (bs="fs") for word is the setting for high-frequency group, for low- frequency group, we directly model the word effect using bs="re" as most words in that category have only one or two tokens. 4 Speech source Frequency [+voi] obstruents [-voi] obstruents sonorant LJ speechHigh1,0001,0001,000 FastSpeech 2High1,0001,0001,000 Tacotron 2High1,0001,0001,000 LJ speechLow1,0001,0001,000 FastSpeech 2Low1,0001,0001,000 Tacotron 2Low1,0001,0001,000 Total6,0006,0006,000 Table 2: Token count per onset type across speech source and frequency bands. 2.4. Additional settings for Experiment 2 While the analysis in Experiment 1 focuses on a single-speaker training corpus, the proposed probing framework can be extended to large-scale and multi-speaker datasets. In this study, we further leverage the ‘In-the-Wild’ dataset (Müller et al., 2022), which contains both bona-fide (real) and deepfake (synthetic) audio from 58 public figures. The real-fake contrast in this dataset provides a natural baseline: bona-fide audio reveals authentic segmental f 0 perturba- tion patterns, while deepfake audio serves as a diagnostic probe to evaluate how well synthetic systems approximate these fine-grained phonetic effects. In this setting, we first normalized f 0 values by applying z-scoring within each speaker. Then, we introduced an additional factor smooth for speaker (m=1, k=5) in the overall GAMMs. This two-step procedure enables meaningful comparison of f 0 trajectories across speakers, particularly in a multi-speaker setting. 2.5. GAMM Structure and Implementation bam(zF0 ~ onset_type + s(time, k=5) + s(time, by = onset_type, k = 5) + s(time, speaker, bs=‘fs’, m=1, k=5) + s(time, speaker, by = onset_type, bs = ’fs’, m = 1, k=5) + s(time, vowel height, bs = ’fs’, m = 1, k=5) + s(time, word, bs = ’fs’, m = 1, k=5) + s(consonant, bs = ’re’), data = data, method = ’fREML’, rho = autocorr_m1, AR.start = data$start.event) Above is the model structure we used for modeling onset-induced f 0 perturbation for the multi-speaker dataset (Experiment 2). For the single-speaker analysis in Experiment 1, we use a simpler model without by speaker factor smoothing. Autocorrelation within f 0 trajectories was modeled using an AR 1 structure. The autocorrelation parame- ter ρ was estimated from a preliminary model using the acf_resid() function from the itsadug package (Van Rij et al., 2015), and applied in the final model to account for temporal dependence in residuals. 3. Corpus analysis In our experimental design, we used lexical frequency as a proxy for corpus familiarity. High-frequency words are statistically more likely to have occurred in the LJ Speech training corpus, while low-frequency words, randomly sampled from the COCA corpus, are less likely to have been encountered during training. Although this frequency- based comparison does not establish a direct causal relationship, it provides a useful basis for testing generalization capacity in the absence of explicit seen and unseen labels. In other words, we aim to evaluate whether models can generalize phonetic detail when the training data is unknown. 5 LJ Speech 200 240 280 05101520 time point f0 (Hz) Onset−f0 effect, overall −50 0 50 100 05101520 time point difference [−voi] obs − [+voi] obs −50 0 50 100 05101520 time point difference [−voi] obs − sonorant −50 0 50 100 05101520 time point difference [+voi] obs − sonorant FastSpeech 2 200 240 280 05101520 time point f0 (Hz) −50 0 50 100 05101520 time point difference −50 0 50 100 05101520 time point difference −50 0 50 100 05101520 time point difference Tacotron 2 200 240 280 05101520 time point f0 (Hz) −50 0 50 100 05101520 time point difference −50 0 50 100 05101520 time point difference −50 0 50 100 05101520 time point difference onset type [−voi] obstruent [+voi] obstruent sonorant significant TRUE FALSE Figure 1: GAMM smooths of f 0 trajectories for onset types (leftmost column) and pairwise difference smooths (right three columns) for all high- frequency speech sources. Shaded areas indicate significant intervals. However, it is still important to confirm the validity of this frequency-based assumption, we performed a corpus- level analysis comparing the word list used in our study against the LJ Speech training transcripts. We found 14,127 unique words in the training corpus, and the majority of the high-frequency items were indeed found in the training corpus (70.69% overlap), whereas most of the low-frequency items were absent (31.74% overlap). This supports our hypothesis that lexical frequency can serve as a reasonable stand-in for training exposure. By examining how f 0 per- turbation patterns differ between these two lexical strata, we are able to assess whether TTS systems rely primarily on memorized surface forms or if they internalize segmental-prosodic relations that support generalization to unfamiliar material. To further testify our memory-versus-generalization hypothesis, we conducted a direct comparison between words that were explicitly seen versus unseen in the training corpus. The results of this analysis are reported in fig.3 below. 4. Analysis and results (Experiment 1) Figures 1 and 2 illustrate the f 0 perturbation patterns for high- and low-frequency words, respectively. In these plots, the high-frequency group consists of words that are highly likely to have appeared in the LJ Speech training corpus, whereas the low-frequency group comprises words that were likely absent from the training data. The low- frequency set thus serves as a diagnostic probe for assessing whether TTS models can generalize f 0 perturbation patterns beyond familiar lexical items. 6 LJ Speech 200 240 280 05101520 time point f0 (Hz) Onset−f0 effect, overall −50 0 50 100 05101520 time point difference [−voi] obs − [+voi] obs −50 0 50 100 05101520 time point difference [−voi] obs − sonorant −50 0 50 100 05101520 time point difference [+voi] obs − sonorant FastSpeech 2 200 240 280 05101520 time point f0 (Hz) −50 0 50 100 05101520 time point difference −50 0 50 100 05101520 time point difference −50 0 50 100 05101520 time point difference Tacotron 2 200 240 280 05101520 time point f0 (Hz) −50 0 50 100 05101520 time point difference −50 0 50 100 05101520 time point difference −50 0 50 100 05101520 time point difference onset type [−voi] obstruent [+voi] obstruent sonorant significant TRUE FALSE Figure 2: GAMM smooths of f 0 trajectories for onset types (leftmost column) and pairwise difference smooths (right three columns) for all low- frequency speech sources. Shaded areas indicate significant intervals. In the high-frequency condition, both the natural recordings (LJ Speech) and the synthetic outputs produced by the TTS models (FastSpeech 2 and Tacotron 2) exhibit similar f 0 perturbation patterns. Specifically, vowels following voiceless obstruents show a sharp initial f 0 peak, while those following voiced obstruents and sonorants exhibit relatively stable f 0 contours. These patterns are consistent with well-attested phonetic findings in English and other languages (Hanson, 2009; Kirby and Ladd, 2016; Yang and Faytak, 2025). Additionally, the f 0 ranges across the three onset types are comparable: 190–220 Hz for voiced obstruents and sonorants, and 240–270 Hz for voiceless obstruents. These observations indicate that both models are capable of reproducing onset-conditioned f 0 patterns when the lexical items are frequent in the training data. In contrast, for low-frequency words, the natural speech recordings maintain the expected f 0 perturbation patterns regardless of lexical frequency. However, the synthetic speech diverges considerably. FastSpeech 2 fails to produce any systematic difference in f 0 contours or height across the three onset types, suggesting a lack of learned segmental- prosodic distinctions in these unseen items. Tacotron 2 shows partial success. Although it appears to capture the elevated f 0 following voiceless obstruents, its f 0 output for vowels following sonorants and voiced obstruents is inaccurate in either overall contour shape or pitch range. 4.1. Direct comparison of the training exposure As discussed above, lexical frequency provides a useful model-agnostic proxy for training exposure, particularly in realistic settings where the exact training data and procedures of TTS systems are not publicly available. However, 7 FastSpeech 2 (seen) 200 240 280 05101520 time point f0 (Hz) Onset−f0 effect, overall −50 0 50 100 05101520 time point diff erence [−voi] obs − [+voi] obs −50 0 50 100 05101520 time point diff erence [−voi] obs − sonorant −50 0 50 100 05101520 time point diff erence [+voi] obs − sonorant FastSpeech 2 (unseen) Lorem ipsum 200 240 280 05101520 time point f0 (Hz) −50 0 50 100 05101520 time point diff erence −50 0 50 100 05101520 time point diff erence −50 0 50 100 05101520 time point diff erence Tacotron 2 (seen) 200 240 280 05101520 time point f0 (Hz) −50 0 50 100 05101520 time point diff erence −50 0 50 100 05101520 time point diff erence −50 0 50 100 05101520 time point diff erence Tacotron 2 (unseen) 200 240 280 05101520 time point f0 (Hz) −50 0 50 100 05101520 time point diff erence −50 0 50 100 05101520 time point diff erence −50 0 50 100 05101520 time point diff erence onset type [−voi] obstruent [+voi] obstruent sonorant significant TRUE FALSE Figure 3: GAMM smooths and difference smooths for z-scored f 0 trajectories (seen vs unseen). we acknowledge that frequency-based operationalizations may introduce noise, as lexical frequency does not always perfectly align with ground-truth training-set membership (seen vs. unseen). To further validate our interpretation and to address this potential source of noise, we additionally conduct a direct comparison based on this ground-truth exposure (seen vs. unseen). Specifically, this experiment tests the effect of training exposure by comparing tokens that are present in the LJ Speech training corpus (“seen”) with those that are absent from it (“unseen”). This analysis allows us to verify that the patterns observed using lexical frequency are not artifacts of the proxy, but reflect genuine differences in model generalization. The same GAMMs specification used in the main analysis was applied here to examine whether onset-conditioned f 0 patterns generalize to completely novel lexical items. Figure 3 presents the resulting smooths and difference smooths for the two conditions. Based on the results, we found that tokens present in the training corpus (“seen”) show robust onset-conditioned separation. Vowels following voiceless obstruents begin at higher f 0 and decrease gradually, with clear and stable temporal differences relative to sonorants and voiced onsets. In contrast, tokens absent from the training corpus (“unseen”) in both FastSpeech 2 and Tacotron 2 exhibit little to no separability, with difference smooths remaining near zero across the vowel. This pattern is consistent with the frequency-stratified analysis and supports a memory- over-abstraction account of how TTS systems realize segmental–prosodic detail. 5. Analysis and results (Experiment 2) This large-scale comparison of onset f 0 perturbation patterns is illustrated in Figures 4 and 5. In the high-frequency group (Fig. 4), both the original and synthetic speech show relatively consistent f 0 trajectories across onset types. 8 Original Speech −1.0 −0.5 0.0 0.5 1.0 05101520 time point zf0 Onset−f0 effect, overall −1.0 −0.5 0.0 0.5 1.0 05101520 time point difference [−voi] obs − [+voi] obs −1.0 −0.5 0.0 0.5 1.0 05101520 time point difference [−voi] obs − sonorant −1.0 −0.5 0.0 0.5 1.0 05101520 time point difference [+voi] obs − sonorant Generated Speech −1.0 −0.5 0.0 0.5 1.0 05101520 time point zf0 −1.0 −0.5 0.0 0.5 1.0 05101520 time point difference −1.0 −0.5 0.0 0.5 1.0 05101520 time point difference −1.0 −0.5 0.0 0.5 1.0 05101520 time point difference onset type [−voi] obstruent [+voi] obstruent sonorant significant TRUE FALSE Figure 4: GAMM smooths and difference smooths for z-scored f 0 trajectories (high frequency group). Voiceless obstruents are followed by a clear rise in f 0 , while sonorants and voiced obstruents exhibit flatter or lower contours. These patterns align with the expectations established from the single-speaker corpus, suggesting that the models can successfully reproduce segmental-prosodic effects when the lexical items are likely familiar. In contrast, the low-frequency group (Fig. 5) presents a different picture. While the original speech continues to display distinct f 0 differences between onset categories, the synthetic speech shows much weaker separation. The trajectories for different onsets become more similar, and the expected rise in f 0 following voiceless obstruents is less consistent. This frequency-based contrast indicates that the ability of TTS models to generate fine-grained prosodic variation diminishes when lexical items are less familiar or unseen during training. While the overall onset-conditioned f 0 patterns remain observable in the In-the-Wild dataset, the effects appear less obvious than those found in the single-speaker LJ Speech corpus. This attenuation is likely due to the large inter- speaker variability inherent in this multi-speaker dataset. In particular, the In-the-Wild corpus includes speakers from diverse linguistic and dialectal backgrounds, and the speech samples cover both formal and informal contexts. These differences can introduce prosodic variation that is not directly related to segmental conditioning, such as differences in intonational phrasing, speaking style, or register. As a result, even with speaker-level modeling, the segmental f 0 effects are slightly less prominent at the group level compared to results in Experiment 1. Another possible contribut- ing factor is that the generative models used for producing speech in the In-the-Wild dataset may have been trained on substantially larger corpora, increasing the likelihood that many test words were already encountered during train- ing. This greater overlap between training and test material could enable the models to reproduce onset-conditioned patterns more consistently from memory, thereby reducing the contrast between high- and low-frequency items and making the overall group-level pattern appear less pronounced. These observations from the In-the-Wild dataset underscore a broader issue that extends beyond the specific sys- tems and datasets examined here. Although this study has examined Tacotron 2, FastSpeech 2, and real versus fake (or natural versus synthetic) speech in an external deepfake dataset, the intention is not to comprehensively bench- mark all available TTS models. Instead, our objective is to illustrate a broader theoretical point. When a model lacks mechanisms for encoding segmental–prosodic relationships in an explicit and structured way, it tends to fail at reproducing fine-grained phonetic effects such as consonant-induced f 0 perturbation. The absence of these patterns 9 Original Speech −1.0 −0.5 0.0 0.5 1.0 05101520 time point zf0 Onset−f0 effect, overall −1.0 −0.5 0.0 0.5 1.0 05101520 time point difference [−voi] obs − [+voi] obs −1.0 −0.5 0.0 0.5 1.0 05101520 time point difference [−voi] obs − sonorant −1.0 −0.5 0.0 0.5 1.0 05101520 time point difference [+voi] obs − sonorant Generated Speech −1.0 −0.5 0.0 0.5 1.0 05101520 time point zf0 −1.0 −0.5 0.0 0.5 1.0 05101520 time point difference −1.0 −0.5 0.0 0.5 1.0 05101520 time point difference −1.0 −0.5 0.0 0.5 1.0 05101520 time point difference onset type [−voi] obstruent [+voi] obstruent sonorant significant TRUE FALSE Figure 5: GAMM smooths and difference smooths for z-scored f 0 trajectories (low frequency group). in low-frequency or unseen items, as observed in our experiments, suggests that this limitation is not specific to any one architecture. Rather, it reflects a tendency in the TTS systems examined here to lack sufficient structural guidance for capturing fine-grained articulatory–acoustic interactions. By treating f 0 perturbation as a diagnostic phenomenon, we offer a linguistically informed framework that can be used to evaluate phonetic generalization in future models, regardless of their underlying architecture. 6. Discussion Our findings suggest that the examined neural TTS systems exhibit only limited ability to generalize segmental prosodic patterns beyond familiar lexical items. While both Tacotron 2 and FastSpeech 2 reproduce f 0 perturbation effects for high-frequency words, their performance degrades sharply on low-frequency words. FastSpeech 2 fails to distinguish between onset types altogether, and Tacotron 2 shows incomplete and unstable patterns. This con- trast points to a reliance on surface-form memorization rather than abstract encoding of articulatory and acoustic interactions, raising concerns about the phonetic depth of learned representations in existing TTS architectures. The comparison between Tacotron 2 and FastSpeech 2 also illustrates that architectural differences do affect how segmen- tal–prosodic patterns are realized, though neither model succeeds fully. The slight difference between the generation results of the two models might be due to the potential role of explicit duration and pitch predictors in FastSpeech 2, which might smooth or oversimplify local prosodic dynamics. These results suggest that neither sequential condi- tioning nor parallel decoding, as currently implemented, is sufficient for capturing fine-grained phonetic interactions without explicit structural guidance. From our results, an additional natural question that arises is why, despite f 0 perturbation being a local phonetic effect and neural models being effective at modeling local acoustic correlations, neural TTS systems fail to realize this effect robustly for low-frequency or unseen words. First, we agree that, from a phonetic perspective, consonant- induced f 0 perturbation is a relatively local phenomenon. We also agree that, in principle, some signal related to this phenomenon should be present even for low-frequency items. Our interpretation, however, concerns how this 10 knowledge is represented and under what conditions it can be reliably accessed. The results suggest that segmental- prosodic effects such as f 0 perturbation are largely encoded in a word-specific manner, tied to lexical items that are well represented in the training data. For high-frequency words, repeated exposure allows the model to reproduce stable and separable onset-conditioned f 0 patterns. In contrast, when the model encounters low-frequency or effec- tively nonce words, it appears unable to consistently invoke a word-independent, compositional segmental–prosodic mapping based solely on onset category. When the model encounters low-frequency or effectively unseen words, these items function as nonce words from the model’s perspective, regardless of whether they are well-formed or even common in the human lexicon. That is, the model lacks the previously lexicalized representations (memorized words) associated with these forms and must rely solely on whatever abstract mechanisms it has learned to relate segmental context to acoustic realization. Based on our understanding, when a TTS model encounters a nonce word that it needs to pronounce, the most important question it has to resolve is not how to model consonant-induced f 0 perturbation. Rather, the primary decision concerns where stress should be placed (on which vowel or vowels)? This is fundamentally a lexical-level problem, not a problem of fine-grained segmental interaction. Because stress assignment largely determines the overall pitch contour, the model’s priority is to ensure global prosodic coherence. As a result, to avoid unstable or erratic pitch behavior, the model might choose to neglect some explicit segmental-level interaction mechanisms. In other words, suppressing consonant-specific f 0 perturbation could be a modeling decision of TTS systems. To illustrate the contrast, imagining pronounce a nonce word such as blicket or splone. Although these words are novel, human speakers can pronounce them in a highly systematic way. For example, their realization still obeys well- established phonetic regularities: articulatory constraints on vocal fold tension, aerodynamic requirements for voicing, and other biomechanical properties of the speech which ensure that consonant-induced f 0 perturbation, among other effects, emerges naturally even in completely novel lexical items. In addition, psychological and phonological con- straints allow speakers to generalize learned sound patterns and apply them productively to new forms. Current neural TTS systems, however, lack access to these articulatory and cognitive constraints. They do not possess an explicit representation of vocal fold dynamics, nor do they encode phonetic rules as abstract, compositional mappings that can be flexibly applied to unseen inputs. As a result, when confronted with a nonce word, the model does not “reason” about how segments should interact based on articulatory principles. Instead, it defaults to an overall statistical tendency shaped by its training data, producing a plausible pronunciation at a global level but failing to consistently instantiate fine-grained segmental–prosodic relations such as onset-conditioned f 0 perturbation. From this perspective, the contrast between human and model behavior is not that models cannot pronounce nonce words at all, but that they do so via fundamentally different mechanisms. Human speakers generate novel words by deploying abstract phonetic knowledge grounded in physiological and cognitive constraints, whereas current TTS models rely on lexicalized and distributional patterns that do not generalize reliably once lexical support is removed. Therefore, the observed frequency effect does not contradict the ability of neural models to learn local correlations. Instead, it highlights a limitation in the generalization and accessibility of such correlations across the lexical space. While weak or implicit CF0-related signals may still be present for low-frequency items, they do not surface as stable, robust, and statistically separable output patterns in group-level modeling. Such limitations can be addressed either by substantially increasing the diversity and quantity of training data, or by introducing explicit mechanisms for modeling segmental–prosodic interactions. The former may allow models to learn generalizable patterns through exposure alone, while the latter offers a more principled approach by guiding the model to encode phonologically meaningful dependencies that are otherwise underrepresented in the data. The observed failure of TTS models to generalize f 0 perturbation to low-frequency words has broader implications beyond phonetic evaluation. First, it highlights a potential vulnerability of neural TTS systems to detection: the absence or degradation of segmental-prosodic effects in rare or novel words may serve as a reliable acoustic cue for distinguishing synthetic from natural speech. This offers a promising direction for developing segmental-level detection tools that go beyond global prosodic or spectral statistics and target fine-grained articulatory correlates. Such segment-based detection is both linguistically grounded and difficult for the current TTS model to accurately reproduce or strategically suppress (Yang et al., 2025). Second, the findings raise the possibility of reverse engineering model characteristics based on acoustic output. For example, the presence or absence of specific phonetic detail, such as onset-driven f 0 modulation, may allow external observers to infer properties of the underlying model, including its architecture, the extent of exposure to particular lexical items, or the granularity of prosodic supervision to some extend. Similarly, the asymmetry in f 0 11 realization across lexical frequency strata suggests that the expressive capacity of a TTS model may be constrained by the lexical and phonetic diversity of its training data. In this sense, systematic probing of segmental effects could be used not only to assess generalization, but also to estimate the effective size or coverage of the training corpus. Similar approaches have been used in other domains. For instance, auditing text-generation models by querying whether a user’s texts were part of the training data (Song and Shmatikov, 2019), or using membership inference attacks to estimate training data memorization in masked language models (Mireshghallah et al., 2022). This opens the door to using phonetic probes as a lens for analyzing TTS training data indirectly, which may have implications for data transparency, privacy auditing, and model interpretability. Finally, an important direction for future work could be the perceptual consequences of failing to reproduce fine- grained segmental effects such as consonant-induced f 0 perturbation in synthetic speech. Although TTS quality as measured by mean opinion scores (MOS) has improved greatly with the advent of modern neural architectures, a persistent gap remains relative to natural speech, particularly for small-scale models such as those examined in this study. While MOS captures overall naturalness and listener preference, it is unspecific to the specific acoustic cues that contribute to these global impressions. Previous work in prosody modeling and expressive TTS has demonstrated that richer and more dynamic pitch contours are associated with higher subjective naturalness and expressiveness (Huang et al., 2023). In the broader speech perception literature, dynamic f 0 variation is known to influence intelligibility and listener preference, with flattened or unnatural pitch contours leading to reduced perceived quality in adverse listening conditions (Laures and Bunton, 2003). These findings align with the intuitive view that sublexical pitch patterns contribute to naturalness judgments, even if they are subtle and not directly captured by aggregate metrics such as MOS. On this basis, it is plausible that the inability of neural TTS systems to consistently generalize local phonetic phenomena to low- frequency or unseen items may contribute in part to the overall MOS gap between synthetic and natural speech. Establishing a direct causal link between segmental-level deviations and listener judgments would require targeted perceptual experiments, such as controlled listening tests or systematic parameter manipulations. Future research could also explore architectural biases, targeted training objectives, and pretraining strategies to enhance the phonetic interpretability and generalization of neural TTS systems. In addition, future research could extend the present experiment to other TTS architectures under controlled training settings, such as diffusion-based and flow-based generative models, to assess whether the observed patterns generalize beyond the models examined here. 7. Conclusion This study examined whether neural TTS systems are capable of reproducing segmental-level f 0 perturbation, a phonetic effect arising from local articulatory mechanisms. By comparing Tacotron 2 and FastSpeech 2 against natural speech, we found that both models can replicate onset-driven f 0 patterns in high-frequency words, but fail to generalize these effects to low-frequency items. These findings suggest that TTS systems rely more on lexical memorization rather than abstract phonetic encoding. Moreover, neither autoregressive nor non-autoregressive archi- tectures, as currently implemented, are sufficient for learning fine-grained segmental–prosodic interactions without explicit guidance. Our proposed probing framework offers a linguistically grounded diagnostic tool for evaluating such interactions, with potential applications in TTS detection, model auditing, and corpus analysis. Appendix A. Data availability and software setup All the speech corpus, lexical corpus, and speech models used in this study are publicly available online, and we have provided corresponding references in the experimental setting section for the purpose of reproduction. Appendix B. Acknowledgment We would like to thank the anonymous reviewers for their constructive comments and helpful suggestions. We are also grateful to Matthew Faytak for helpful comments and discussions that later proved valuable in shaping some of the ideas developed in this paper. 12 References Boersma, P., 2007. Praat: doing phonetics by computer. http://w.praat.org/ . Brysbaert, M., New, B., 2009. Moving beyond Ku ˇ cera and Francis: A critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for American English. Behavior research methods 41, 977–990. Chen, Y., 2011. How does phonology guide phonetics in segment–f0 interaction? Journal of Phonetics 39, 612–625. Clements, G.N., Khatiwada, R., 2007. Phonetic realization of contrastively aspirated affricates in Nepali, in: Proceed- ings of the 16th International Congress of Phonetic Sciences, p. 632. Davies, M., 2015. The Corpus of Contemporary American English (COCA), version 2.2. Harvard Dataverse. doi:10.7910/DVN/AMUDUW. one-billion-word balanced corpus (1990–present). Gahl, S., 2008. Time and thyme are not homophones: The effect of lemma frequency on word durations in spontaneous speech. Language 84, 474–496. Gao, J., Kirby, J., 2024. Laryngeal contrast and sound change: The production and perception of plosive voicing and co-intrinsic pitch. Language 100, 124–158. Halle, M., Stevens, K.N., 1971. A note on laryngeal features. Quarterly progress report 101, 198–213. Hanson, H.M., 2009. Effects of obstruent consonants on fundamental frequency at vowel onset in English. The Journal of the Acoustical Society of America 125, 425–441. doi:10.1121/1.3021306. Hoole, P., Honda, K., 2011. Automaticity vs. feature-enhancement in the control of segmental F0, in: Where do phonological features come from? Cognitive, physical and developmental bases of distinctive speech categories. John Benjamins Publishing Company, p. 131–172. Huang, R., Zhang, C., Ren, Y., Zhao, Z., Yu, D., 2023. Prosody-tts: Improving prosody with masked autoencoder and conditional diffusion model for expressive text-to-speech, in: Findings of the Association for Computational Linguistics: ACL 2023, p. 8018–8034. Ito, K., Johnson, L., 2017. The lj speech dataset. https://keithito.com/LJ-Speech-Dataset/. Jacewicz, E., Fox, R.A., 2015. Intrinsic fundamental frequency of vowels is moderated by regional dialect. The Journal of the Acoustical Society of America 138, EL405–EL410. Kim, J., Kong, J., Son, J., 2021. Conditional variational autoencoder with adversarial learning for end-to-end text-to- speech, in: International Conference on Machine Learning, PMLR. p. 5530–5540. Kingston, J., Diehl, R.L., 1994. Phonetic knowledge. Language 70, 419–454. Kirby, J.P., Ladd, D.R., 2016. Effects of obstruent voicing on vowel F0: Evidence from “true voicing” languages. The Journal of the Acoustical Society of America 140, 2400–2411. Kong, J., Kim, J., Bae, J., 2020. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in neural information processing systems 33, 17022–17033. Laures, J.S., Bunton, K., 2003. Perceptual effects of a flattened fundamental frequency at the sentence level under different listening conditions. Journal of communication disorders 36, 449–464. Löfqvist, A., Baer, T., McGarr, N.S., Story, R.S., 1989. The cricothyroid muscle in voicing control. The Journal of the Acoustical Society of America 85, 1314–1321. McAuliffe, M., Socolof, M., Mihuc, S., Wagner, M., Sonderegger, M., 2017a. Montreal forced aligner: Trainable text-speech alignment using kaldi., in: Interspeech, p. 498–502. 13 McAuliffe, M., Sonderegger, M., 2022. English (US) MFA dictionary v2.0.0a. Technical Report. Montreal Forced Aligner Project. McAuliffe, M., Stengel-Eskin, E., Socolof, M., Sonderegger, M., 2017b. Polyglot and Speech Corpus Tools: A System for Representing, Integrating, and Querying Speech Corpora, in: Interspeech 2017, p. 3887–3891. doi:10.21437/Interspeech.2017-1390. Mireshghallah, F., Goyal, K., Uniyal, A., Berg-Kirkpatrick, T., Shokri, R., 2022. Quantifying privacy risks of masked language models using membership inference attacks. arXiv preprint arXiv:2203.03929 . Müller, N.M., Czempin, P., Dieckmann, F., Froghyar, A., Böttinger, K., 2022. Does audio deepfake detection gener- alize? Interspeech . Park, Y., Stepp, C.E., 2019. The effects of stress type, vowel identity, baseline f0, and loudness on the relative fundamental frequency of individuals with healthy voices. Journal of Voice 33, 603–610. Ren, Y., Hu, C., Tan, X., Qin, T., Zhao, S., Zhao, Z., Liu, T.Y., 2020. Fastspeech 2: Fast and high-quality end-to-end text to speech. arXiv preprint arXiv:2006.04558 . Rose, P., Yang, T., 2022. Modelling Interaction between tone and phonation type in the Northern Wu dialect of Jinshan, in: Proc. 18th Int’l Australasian Conf. on Speech Science & Technology, p. 221–225. Schertz, J., Khan, S., 2020. Acoustic cues in production and perception of the four-way stop laryngeal contrast in Hindi and Urdu. Journal of Phonetics 81, 100979. Shen, J., Pang, R., Weiss, R.J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Skerrv-Ryan, R., et al., 2018. Natural tts synthesis by conditioning wavenet on mel spectrogram predictions, in: 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP), IEEE. p. 4779–4783. Solé, M.J., 2018. Articulatory adjustments in initial voiced stops in Spanish, French and English. Journal of Phonetics 66, 217–241. Song, C., Shmatikov, V., 2019. Auditing data provenance in text-generation models, in: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, p. 196–206. Sóskuthy, M., 2017. Generalised additive mixed models for dynamic analysis in linguistics: A practical introduction. arXiv preprint arXiv:1703.05339 doi:10.48550/arXiv.1703.05339. Ting, C., Sonderegger, M., Clayards, M., McAuliffe, M., 2025. The crosslinguistic distribution of vowel and consonant intrinsic F0 effects. Language 101, 1–36. Van Den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., Kavukcuoglu, K., et al., 2016. Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499 12, 1. Van Hoof, S., Verhoeven, J., 2011. Intrinsic vowel F0, the size of vowel inventories and second language acquisition. Journal of Phonetics 39, 168–177. Van Rij, J., Wieling, M., Baayen, R.H., van Rijn, D., 2015. itsadug: Interpreting time series and autocorrelated data using gamms . Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R.J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., et al., 2017. Tacotron: Towards end-to-end speech synthesis. arXiv preprint arXiv:1703.10135 . Whalen, D.H., Levitt, A.G., 1995. The universality of intrinsic F0 of vowels. Journal of phonetics 23, 349–366. Whalen, D.H., Levitt, A.G., Hsiao, P.L., Smorodinsky, I., 1995. Intrinsic F0 of vowels in the babbling of 6-, 9- , and 12-month-old french-and english-learning infants. The Journal of the Acoustical Society of America 97, 2533–2539. 14 Wieling, M., 2018. Analyzing dynamic phonetic data using generalized additive mixed modeling: A tutorial focusing on articulatory differences between l1 and l2 speakers of English. Journal of Phonetics 70, 86–116. Williams, D., Escudero, P., 2014. A cross-dialectal acoustic comparison of vowels in Northern and Southern british english. The Journal of the acoustical society of America 136, 2751–2761. Wood, S.N., 2017. Generalized additive models: an introduction with R. chapman and hall/CRC. Xu, C.X., Xu, Y., 2003. Effects of consonant aspiration on Mandarin tones. Journal of the International Phonetic Association 33, 165–181. Xu, Y., Xu, A., 2021. Consonantal F0 perturbation in American English involves multiple mechanisms. The Journal of the Acoustical Society of America 149, 2877–2895. Yang, T., Faytak, M., 2025. Onset-tone interaction in Mundabli. Proceedings of the Linguistic Society of America 10, 5895–5895. Yang,T.,Sun,C.,Lyu,S.,Rose,P.,2025.Forensicdeepfakeaudiodetec- tionusingsegmentalspeechfeatures.ForensicScienceInternational379,112768. URL: https://w.sciencedirect.com/science/article/pii/S0379073825004128, doi:https://doi.org/10.1016/j.forsciint.2025.112768. Ye, Z., Huang, R., Ren, Y., Jiang, Z., Liu, J., He, J., Yin, X., Zhao, Z., 2023. Clapspeech: Learning prosody from text context with contrastive language-audio pre-training. arXiv preprint arXiv:2305.10763 . Zhang, M., 2020. Effect of consonants on onset F0: Evidence from Kansai Japanese. The Journal of the Acoustical Society of America 148, 2472–2472. 15