Paper deep dive
Learning-free L2-Accented Speech Generation using Phonological Rules
Thanathai Lertpetchpun, Yoonjeong Lee, Jihwan Lee, Tiantian Feng, Dani Byrd, Shrikanth Narayanan
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/13/2026, 12:35:13 AM
Summary
The paper introduces a learning-free framework for generating L2-accented English speech (specifically Spanish and Indian accents) by applying phonological transformation rules to phoneme sequences before synthesis with a pretrained multilingual Text-to-Speech (TTS) model. This approach avoids the need for large-scale accented training datasets while enabling fine-grained phoneme-level control and maintaining speech naturalness.
Entities (5)
Relation Signals (3)
Thanathai Lertpetchpun â authored â Learning-free L2-Accented Speech Generation using Phonological Rules
confidence 100% ¡ Learning-free L2-Accented Speech Generation using Phonological Rules Thanathai Lertpetchpun
Kokoro-82M â usedfor â Accented Speech Generation
confidence 95% ¡ We use a pretrained multilingual TTS model, Kokoro-82M v0.19, to generate all speech samples.
Phonological Rules â transforms â American English
confidence 90% ¡ The rules are applied to phoneme sequences to transform accent at the phoneme level
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Accent plays a crucial role in speaker identity and inclusivity in speech technologies. Existing accented text-to-speech (TTS) systems either require large-scale accented datasets or lack fine-grained phoneme-level controllability. We propose a accented TTS framework that combines phonological rules with a multilingual TTS model. The rules are applied to phoneme sequences to transform accent at the phoneme level while preserving intelligibility. The method requires no accented training data and enables explicit phoneme-level accent manipulation. We design rule sets for Spanish- and Indian-accented English, modeling systematic differences in consonants, vowels, and syllable structure arising from phonotactic constraints. We analyze the trade-off between phoneme-level duration alignment and accent as realized in speech timing. Experimental results demonstrate effective accent shift while maintaining speech quality.
Tags
Links
- Source: https://arxiv.org/abs/2603.07550v1
- Canonical: https://arxiv.org/abs/2603.07550v1
Trouble viewing inline? Open PDF directly â
Full Text
28,609 characters extracted from source content.
Expand or collapse full text
Learning-free L2-Accented Speech Generation using Phonological Rules Thanathai Lertpetchpun ID 1 , Yoonjeong Lee ID 1 , Jihwan Lee ID 1 , Tiantian Feng ID 1 , Dani Byrd ID 2 , Shrikanth Narayanan ID 1,2 1 Signal Analysis and Interpretation Lab, University of Southern California, USA 2 Department of Linguistics, University of Southern California lertpetc@usc.edu Abstract Accent plays a crucial role in speaker identity and inclu- sivity in speech technologies. Existing accented text-to-speech (TTS) systems either require large-scale accented datasets or lack fine-grained phoneme-level controllability. We propose a accented TTS framework that combines phonological rules with a multilingual TTS model. The rules are applied to phoneme sequences to transform accent at the phoneme level while pre- serving intelligibility. The method requires no accented training data and enables explicit phoneme-level accent manipulation. We design rule sets for Spanish- and Indian-accented English, modeling systematic differences in consonants, vowels, and syl- lable structure arising from phonotactic constraints. We analyze the trade-off between phoneme-level duration alignment and ac- cent as realized in speech timing. Experimental results demon- strate effective accent shift while maintaining speech quality. Index Terms: Accented Text-To-Speech Model, Speech Syn- thesis, Phonological Rules, Speech Generation 1. Introduction Despite the status of English as the global lingua franca, its linguistic landscape is remarkably decentralized. Current es- timates indicate that non-native (L2) speakers outnumber na- tive (L1) speakers by a ratio of three to one [1, 2]. This demo- graphic reality is mirrored by a vast spectrum of phonetic di- versity, ranging from regional L1 varietiesâsuch as American, British, and Singaporean English [3]âto L2 varieties charac- terized by the phonological influence and distinct phonotactic constraints of a variety of native (L1) languages [4, 5]. However, existing Text-to-Speech (TTS) systems have his- torically focused on a narrow subset of mainstream accents (e.g., North American or British English), often failing to repre- sent the authentic speech patterns of the global majority [6â9]. This gap in development presents several challenges. Such TTS models often treat accented speech as deviation rather than as systematic variations, leading to poor synthesis quality for di- verse user bases. For L2 listeners, synthetic speech that does not align with familiar phonological structures can increase pro- cessing effort and reduce overall intelligibility. To address these disparities, there is a need for accented TTS systems capable of modeling the nuanced phonetic signatures of both regional and L2 English. Accented text-to-speech (TTS) models aim to control the accent of synthesized speech.A common approach is to fine-tune a pretrained TTS model on accented datasets [10, 11]. However, collecting large-scale, high-quality L2-accented speech corpora is costly and time-consuming.To circum- vent this limitation, [12] propose using a large language model (LLM) to transliterate input text into the orthography of a tar- get language, followed by synthesis with a multilingual TTS model. While this approach removes the dependency on ac- cented datasets, it typically yields a fixed accent style and lacks fine-grained phonetic controllability. More recently, [13] demonstrate that phonological rule- based transformations can enable controllable accent strength through systematic phoneme modifications. In particular, they propose a set of US-to-UK phonological transformation rules, achieving fine-grained control at the phoneme level and effec- tive accent manipulation between American and British En- glish. Nevertheless, this method is limited to accent variation within the same language and has not been extended to cross- lingual accent transfer. Building on these insights, we explore a multistage pipeline grounded in multilingual TTS modeling. Many multilingual TTS systems are conditioned on both speaker embeddings and phoneme sequences [14â17].Speaker embeddings capture language and accent characteristics [18], while phoneme se- quences provide a largely speaker-agnostic pronunciation rep- resentation. This architecture enables accent transformation by modifying phoneme inputs without retraining the TTS model. Specifically, we first apply a phonological rule system to con- vert American English phoneme sequences into target-accent variants (e.g., Spanish-accented or Indian-accented English). The modified phoneme sequences are then synthesized using a pretrained multilingual TTS model. Leveraging a multilingual backbone allows the system to draw upon its cross-lingual prior knowledge, producing speech that reflects not only segmental pronunciation alternations but also suprasegmental characteris- tics such as rhythm and intonation. In summary, our contributions are as follows: 1) We pro- pose a phonology-driven framework for generating accented speech using a pretrained multilingual TTS model, requiring no accented training data. 2) We enable fine-grained control over accent strength through selective phoneme-level transfor- mations implemented as a lightweight preprocessing step with no additional model training. 3) We analyze rhythmic variations influenced by speakersâ native languages and experimentally validate the effectiveness of the proposed approach. Speech samples are publicly available 1 . 2. Methods 2.1. Phonological Rules We initiate this line of research by focusing on two target ac- cents: Spanish-accented English (SP) and Indian-accented En- glish (IN). The phonological transformation rules are derived 1 For review, refer to the supplementary materials. Full source code will be released upon acceptance. arXiv:2603.07550v1 [cs.CL] 8 Mar 2026 Table 1: Phonological transformation rules that demonstrates shifts from US English to Spanish- and Indic-accented English. Spanish RuleUS IPASpanish IPA 1. Initial Consonant Substitution/v T D z Ă//b s d s j/ 2. Spanish Rhoticity/Ă´//r R/ 3. Epenthesis (s-clusters)/sp st sk//esp est esk/ 4. Final Consonant Devoicing/b d g z Ă//p t k s Ă/ 5. Vowel Simplification/I U @ A 2 E 3 O//i u a a a e e o a/ 6. Monophthongization and Schwa/eI oU @//e o a/ Indian RuleUS IPAIndian IPA 1. Retroflexion of Stops and R/t d Ă´//Ăş ĂŁ Ă/ 2. Dentalization of Fricatives/T D ĂŹ//tâ dâ l/ 3. Consonant Substitutions/v Z//w z/ 4. Vowel Simplification/I U ĂŚ 2! E 3 O//i u a @ a e e o/ 5. Monophthongization and Schwa/eI oU @//e o a/ Table 2: Examples of applying phonological rules. The first col- umn indicates which accented IPA they are and the third column indicates which rules were applied for the transformation IPA ExamplesRules This very tall teacher closed the big park US/DIs vEĂ´i tOl tiĂ@Ă´ klOzd D@ bIg pAĂ´k./ SP/dIs bERi dol tiĂ@R glozd d@ bik baĂ´k./1,2,4,5,6 IN/dâIs wERi Ăşol ĂşiĂ@R klozĂŁ dâ@ big paĂ´k./1,2,3,4,5 The little button was very broken. US/D@ lIR@l b2tn w2z vEĂ´i bĂ´Ok@n./ SP/d@ liRal batn was bERi bĂ´ok@n./2,5,6,7 IN/dâ@ lIR@l b@tn w@z wERi bĂ´ok@n./1,2,3,4,5 Kate found three big red stones. US/keIt faUd TĂ´i bIg Ă´Ed stOnz./ SP/get faUt sĂ´i bik Ă´ed estOns./1,3,4,5,6 IN/keIt faUd tâĂ´i big REd sĂşOnz./1,3,4,5 from (1) linguistically documented phonotactic and phonemic properties of the respective L1 systems and (2) systematic phonological differences between American English and these L1s. Each rule models salient and well-documented cross- lingual phonological interactions [19â21]. Table 1 summarizes the currently implemented phonologi- cal transformation rules. For each accent, phonemes in Ameri- can English (middle column) are mapped to their corresponding realizations in the target accent (right column). For example, the word three, /TĂ´i/ is realized as [sĂ´i] under the Spanish accent by applying Rule 1 (Initial Consonant Substitution), which maps /T/ to /s/. We exclude potential rules for some minor changes in which the accent difference is handled more simply by the speaker embedding. Table 2 provides sentence-level examples illustrating the application of these transformations. The first row presents the American IPA transcription, while the second and third rows show the transformed Spanish- and Indian-accented outputs, re- spectively. The final column indicates which rules were applied to generate each accented form. The transformations are ap- plied deterministically based on the defined rule set. 2.2. Accented Speech Generation We leverage an existing multilingual TTS model to generate ac- cented speech. As illustrated in Fig. 1, the multilingual TTS model takes two inputs: a speaker embedding and a phoneme sequence. In multilingual settings, the speaker embedding typi- cally captures speaker identity and language and accent charac- teristics. In our approach, we exploit this implicit language encod- Figure 1: Synthesis Pipeline. We follow the synthesis and eval- uation pipeline described in [13]. The key difference is that we condition the TTS model on a speaker embedding from the tar- get language. In addition, we explicitly control whether dura- tion alignment is applied between the American (US) phoneme sequence and the transformed target-accent phoneme sequence. ing in the speaker embedding to synthesize accented English. Specifically, we condition the TTS model on a speaker em- bedding corresponding to a target L1 (e.g., Spanish or In- dian), while the phoneme sequence represents English content modified by accent-specific phonological rules. The phoneme transformation rules convert American English phonemes into their accented counterparts (e.g., Spanish-accented or Indian- accented English). By combining a target-language speaker embedding with systematically transformed English phonemes, the model generates English speech exhibiting the desired L1- accent characteristics. 2.3. Rhythmic Differences We further investigate cross-linguistic rhythmic differences and their influence on L2-accented speech. Since L2 pronunciation is strongly influenced by speakersâ L1 prosodic system, rhyth- mic or timing transfer effects are expected to exist in accented English. For instance, the major Indian language, Hindi, is commonly characterized as a syllable-timed language, in which sequential syllables tend to exhibit relatively regular durations compared to stress-timed languages such as English, which have more oscillatory segment durations (and vowel reduction) for syllables in sequence [22]. This more even temporal dis- tribution can influence how Indian speakers realize rhythm in English. As a result, Indian-accented English often exhibits less durational contrast between stressed and unstressed sylla- bles than L1 English. To examine the impact of rhythmic transfer, we manipulate duration modeling in the TTS system. Specifically, we compare conditions in which accent-specific duration patterns are pre- served versus conditions in which durations are force-aligned to American English timing. As our framework allows explicit control over duration prediction, we can impose American En- glish durations onto Spanish- or Indian-accented phoneme se- quences. This design enables us to isolate and analyze the con- tribution of rhythmic differences to perceived accent character- istics. Table 3: Effect of duration alignment on accentedness. Accent classification probabilities are reported for UK, Spanish (SP), and Indian (IN) conditions with and without phoneme-level du- ration alignment. Probability UKSPIN UK With Alignment80.281.230.17 UK Without Alignment 81.121.060.7 SP With Alignment6.50 51.593.28 SP Without Alignment3.950.862.73 IN With Alignment0.826.0886.4 IN Without Alignment0.562.88 93.1 3. Experiments 3.1. Datasets and Experimental Setup We use a pretrained multilingual TTS model, Kokoro-82M v0.19 2 , to generate all speech samples. The model provides 20 American, 8 British, 3 Spanish, and 4 Hindi speaker embed- dings. To ensure consistency across experiments, we select one representative embedding for each language group: af heart (UK), bmfable (US), efdora (Spanish), and hfalpha (Indian). For the main experiments (Tables 4 and 3), we use transcripts from the LibriTTS-R [23] train-clean-100 split, comprising ap- proximately 33k utterances. For the rule ablation study (Ta- ble 5), we use transcripts from the LibriTTS-R test-clean split, containing approximately 4.8k utterances. 3.2. Evaluation Metrics Following [13], we evaluate the proposed phonological rules from two perspectives: accent strength and synthesis quality. Each aspect is evaluated by both objective and subjective met- rics. Accent Strength To evaluate whether the synthesized speech reflects the intended target accent, we employ the narrow ver- sion of Vox-Profile [24], which outputs posterior probabilities over 16 accent categories. Since Vox-Profile does not provide explicit labels for Spanish-accented English or Indian-accented English, we use Romance and South Asia as proxy labels, re- spectively. We additionally compute cosine similarity in the Vox-Profile accent embedding space. We report the average co- sine similarity as a measure of accent proximity in the learned embedding space. Synthesis Quality Naturalness is measured using UTMOS [25], an automatic predictor of mean opinion score (MOS). The predicted score ranges from 1 to 5, corresponding to least nat- ural to human-level naturalness. Intelligibility is evaluated us- ing Whisper-medium [26] as an Automatic Speech Recognition (ASR) system. We report Word Error Rate (WER) and Charac- ter Error Rate (CER) computed between the ASR transcription and the reference text. Subjective Evaluation Along with the objective measures above, we also conduct subjective evaluation on perceived ac- cents, accent strength, and naturalness by human listeners. 14 human evaluators were recruited, and they are either native or fluent speakers of English. For non-native fluent speakers of English, they currently reside in the US, and their L1 back- grounds include various Asian and European languages, such as Korean, Chinese, Thai, Japanese, Indian English, Norwe- 2 https://huggingface.co/hexgrad/Kokoro-82M. Table 4: Overall effects of phonological rules on accentedness and speech quality. Accent probability and embedding similar- ity are reported for US and target accents, along with UTMOS, WER, and CER. ProbabilitySimilaritySpeech Quality US Target US Target UTMOS WER CER US Spk Emb. 73.8-0.71-4.383.422.83 +SP rules25.97-0.50-4.3924.91 13.56 +IN rules16.85-0.42-4.3832.66 18.32 SP Spk Emb.26.6 23.70.73 0.473.737.995.64 +SP rules2.97 51.590.38 0.603.8713.137.79 IN Spk Emb.1.89 58.860.28 0.744.258.355.68 +IN rules0.27 86.4-0.008 0.754.1624.15 14.78 gian, and Danish. A total of 70 samples are evaluated. Lis- teners were first asked to choose what the perceived accent is, and then followed that by rating how prominent the accent is on the following scale: [1:not at all prominent, 2:slightly promi- nent, 3:moderately prominent, 4:quite prominent, 5:extremely prominent]. Lastly, listeners were asked to rate the sampleâs resemblance to human speech the same 1-5 scale. 4. Results and Discussion We evaluate our approach under three experimental setups: 1) Overall effectiveness of phonological rules: We measure ac- centedness by reporting classification probabilities and embed- ding similarities for American (US), Spanish (SP), and Indian (IN) accents. We additionally report UTMOS, WER, and CER to evaluate naturalness and intelligibility. 2) Effect of dura- tion alignment: We compare results with and without duration alignment for each target accent after applying the phonolog- ical rules, in order to analyze the impact of rhythmic control. 3) Rule ablation analysis: We incrementally add individual phonological rules for each accent and measure their separate contribution to accentedness and overall performance. 4.1. Phonological Rules Table 4 demonstrates the effectiveness of the proposed phono- logical rules. In the first section, the baseline condition applies only the speaker embedding without phonological modification, while the second and third columns correspond to applying the Spanish and Indian phonological rules, respectively. When the rules are applied, the probability and embedding similarity of an American (US) accent decrease, indicating a successful shift away from an American accent. This shift is accompanied by an increase in WER and CER. It is important to recognizer that WER reflects both intel- ligibility and accentedness. Since most ASR systems, includ- ing Whisper, are predominantly trained on American English, they may inherently produce higher error rates for accented speech [27]. For example, one of our Spanish-accented rules transforms the word think (/Tink/) to the accented pronuncia- tion sink (/sink/) so as to appropriately alter the American ac- cent toward the Spanish-accented English. An ASR system may mark this as an error, even though the pronunciation change is the intended and appropriately accented outcome of the phono- logical rule. Therefore, WER and CER should be interpreted with caution, as they may partially capture accent-induced pho- netic variation rather than any intelligibility degradation per se; of course it is possible that accent can in real-world contexts de- grade intelligibility variably for listeners, particularly for those with less experience with the target accent or for accents that utilize phones that are more distinct from the listenerâs own. Table 5: Effect of individual phonological rules on accented- ness. Accent probability and embedding similarity are shown as each rule is added to the Spanish (SP) and Indian (IN) speaker embeddings. Accent ProbAccent Sim USâ SPâ USâ SPâ SP Spk Emb32.4622.860.700.32 + Rule 129.8022.200.700.37 + Rule 221.1323.000.660.32 + Rule 330.5022.780.690.32 + Rule 428.4725.220.690.35 + Rule 59.82 25.72 0.59 0.44 + Rule 624.2520.120.650.33 + All Rules2.26 50.33 0.43 0.58 USâ INâ USâ INâ IN Spk Emb1.6570.680.250.78 + Rule 10.18 94.63 0.06 0.83 + Rule 21.1873.570.220.79 + Rule 31.3072.930.220.79 + Rule 41.0762.640.120.73 + Rule 51.0479.180.200.79 + All Rules0.16 92.45 -0.03 0.78 In the second and third sections of the table, we use the Spanish and Indian speaker embeddings as baselines and com- pare them with conditions where the corresponding phonolog- ical rules are additionally applied. In both cases, applying the rules leads to higher target-accent probability and embedding similarity, confirming the effectiveness of the proposed trans- formations in producing the desired accented English. Across all conditions, UTMOS scores remain stable, sug- gesting that the phonological modifications do not degrade per- ceptual naturalness and that the synthesized speech maintains consistent quality. 4.2. Rhythmic Variation Effect Table 3 illustrates the effect of rhythmic control through phoneme-level duration alignment. We compare conditions with and without phoneme-level alignment to examine how tim- ing influences accentedness. For the UK and Indian (IN) con- ditions, removing phoneme-level alignment results in higher target-accent probability, suggesting that preserving accent- specific phone timing patterns strengthens accent perception. For the Spanish condition, although the target-accent probabil- ity does not increase, the probabilities of US and IN accents de- crease. This indicates reduced similarity to non-target accents and suggests that timing variation contributes to accent differen- tiation. Overall, these results highlight the role of timing-based rhythmic patterns in shaping perceived accent characteristics. 4.3. Effect of Each Phonological Rule Table 5 demonstrates the effectiveness of each rule on trans- forming the accent. The rules are ordered in the same order as in 1. We first apply the Spanish or Indian speaker embedding and apply the rules one by one to see their separable effects. The results show that rule 5 (Vowel Simplification) is the most influential in transforming the American accent to Spanish ac- cent. While in Indian accent, rule 1 (Retroflexion of Stops and R) is the most prominent feature. Clearly, employing all rules together is the most effective way to transforming the accent. Table 6: Human evaluation of accent perception. We report lis- tener judgments of accent accuracy, perceived accent strength (for correctly identified samples), and naturalness. Accuracy Strength Naturalness US Spk Emb 1.0004.044.09 + SP rules0.3862.592.81 + IN rules0.2292.882.73 SP Spk Emb0.0712.203.14 + SP rules0.7573.082.94 IN Spk Emb0.7863.313.47 + IN rules0.7573.943.14 4.4. Subjective Evaluation As shown in Table 6 and Fig. 2, applying phonological rules to the US speaker embedding noticeably shifts listenersâ per- ception of the accent, reducing its identification as American. When using only the Spanish speaker embedding, most listen- ers still perceived the speech as American, suggesting that the embedding alone is insufficient to produce a strong Spanish accent. In contrast, incorporating Spanish-specific phonolog- ical rules substantially improves both accent identification ac- curacy and perceived accent strength. For the Indian accent, the speaker embedding already introduces distinguishable ac- cent cues; however, applying additional phonological rules in- creases perceived strength while also introducing some confu- sion between Spanish and Indian accents. Importantly, natural- ness ratings remain around 3 (âModerately Naturalâ) across all conditions, indicating that accent manipulation does not signif- icantly degrade overall perceptual quality. Figure 2: Human accent confusion matrix. Rows denote syn- thesis conditions (speaker embeddings with and without phono- logical rules), and columns represent listener-labeled accents. 5. Conclusion In this work, we propose phonological transformation rules that map American phonemes to Spanish- and Indian-accented vari- ants. By integrating these rules into a pretrained multilingual TTS model, we generate accented speech without requiring ac- cented training data. We further analyze rhythmic variation, phonological rule strength, and human evaluations confirm the perceptual effectiveness of our approach. 6. Acknowledgement This work was supported by the Office of the Director of Na- tional Intelligence (ODNI), Intelligence Advanced Research Projects Activity (IARPA), via the ARTS Program under con- tract D2023-2308110001. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies, either expressed or implied, of ODNI, IARPA, or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for governmental purposes notwithstanding any copyright an- notation therein. 7. Generative AI Use Disclosure Generative AI tools were employed to assist with code develop- ment, manuscript editing, and language refinement. The foun- dational research concepts, problem formulation, methodology, and data analysis were developed exclusively by the authors. The authors retain full accountability for the integrity and accu- racy of the final manuscript. 8. References [1] D. Crystal, English as a Global Language, 2nd ed.Cambridge University Press, 2003. [2] Ethnologue, âEnglish language profile,â https://w.ethnologue. com/language/eng, 2024, accessed: 2026-02-23. [3] J. C. Wells, Accents of English.Cambridge University Press, 1982. [4] B. B. Kachru, The Alchemy of English: The Spread, Functions, and Models of Non-Native Englishes. University of Illinois Press, 1990. [5] E. W. Schneider, Postcolonial English: Varieties around the World. Cambridge University Press, 2007. [6] E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. G Ě olge, and M. A. Ponti, âYourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,â in International confer- ence on machine learning. PMLR, 2022, p. 2709â2720. [7] J. Kim, J. Kong, and J. Son, âConditional variational autoencoder with adversarial learning for end-to-end text-to-speech,â in Inter- national conference on machine learning.PMLR, 2021, p. 5530â5540. [8] S. Chen, C. Wang, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li et al., âNeural codec language models are zero-shot text to speech synthesizers,â IEEE Transactions on Audio, Speech and Language Processing, vol. 33, p. 705â718, 2025. [9] S. Chen, S. Liu, L. Zhou, Y. Liu, X. Tan, J. Li, S. Zhao, Y. Qian, and F. Wei, âVall-e 2: Neural codec language models are hu- man parity zero-shot text to speech synthesizers,â arXiv preprint arXiv:2406.05370, 2024. [10] R. Liu, B. Sisman, G. Gao, and H. Li, âControllable accented text- to-speech synthesis with fine and coarse-grained intensity render- ing,â IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, p. 2188â2201, 2024. [11] X. Zhou, M. Zhang, Y. Zhou, Z. Wu, and H. Li, âAccented text- to-speech synthesis with limited data,â IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, p. 1699â 1711, 2024. [12] S. Inoue, S. Wang, W. Wang, P. Zhu, M. Bi, and H. Li, âMacst: Multi-accent speech synthesis via text transliteration for accent conversion,â in ICASSP 2025-2025 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, p. 1â5. [13] T. Lertpetchpun, Y. Lee, T. Trachu, J. Lee, T. Feng, D. Byrd, and S. Narayanan, âQuantifying speaker embedding phonologi- cal rule interactions in accented speech synthesis,â arXiv preprint arXiv:2601.14417, 2026. [14] Z. Zhang, L. Zhou, C. Wang, S. Chen, Y. Wu, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li et al., âSpeak foreign languages with your own voice: Cross-lingual neural codec language modeling,â arXiv preprint arXiv:2303.03926, 2023. [15] C. Tran, C. M. Luong, and S. Sakti, âSTEN-TTS: Improving Zero-shot Cross-Lingual Transfer for Multi-Lingual TTS with Style-Enhanced Normalization Diffusion Framework,â in Inter- speech 2023, 2023, p. 4464â4468. [16] C. Lu, X. Wen, L. Song, and J. Oh, âRobust neural codec language modeling with phoneme position prediction for zero-shot tts,â in Proc. Interspeech 2025, 2025, p. 2475â2479. [17] K. Fujita, A. Ando, and Y. Ijima, âSpeech rhythm-based speaker embeddings extraction from phonemes and phoneme duration for multi-speaker speech synthesis,â IEICE TRANSACTIONS on In- formation and Systems, vol. 107, no. 1, p. 93â104, 2024. [18] S. Liu, Y. Guo, C. Du, X. Chen, and K. Yu, âDse-tts: Dual speaker embedding for cross-lingual text-to-speech,â INTER- SPEECH 2023, p. 616â620, 2023. [19] L. Fabiano-Smith and B. A. Goldstein, âPhonological acquisi- tion in bilingual spanishâenglish speaking children,â Journal of speech, language, and hearing research, vol. 53, no. 1, p. 160â 178, 2010. [20] P. Patel, N. Chatterjee Singh, and M. Torppa, âUnderstanding the role of cross-language transfer of phonological awareness in emergent hindiâenglish biliteracy acquisition,â Reading and Writ- ing, vol. 37, no. 4, p. 887â920, 2024. [21] A. Gottardo, âThe relationship between language and reading skills in bilingual spanish-english speakers,â Topics in language disorders, vol. 22, no. 5, p. 46â70, 2002. [22] H. Sirsa and M. A. Redford, âThe effects of native language on indian english sounds and timing patterns,â Journal of phonetics, vol. 41, no. 6, p. 393â406, 2013. [23] Y. Koizumi, H. Zen, S. Karita, Y. Ding, K. Yatabe, N. Morioka, M. Bacchiani, Y. Zhang, W. Han, and A. Bapna, âLibritts-r: A re- stored multi-speaker text-to-speech corpus,â in Interspeech 2023, 2023, p. 5496â5500. [24] T. Feng, J. Lee, A. Xu, Y. Lee, T. Lertpetchpun, X. Shi, H. Wang, T. Thebaud, L. Moro-Velazquez, D. Byrd et al., âVox-profile: A speech foundation model benchmark for characterizing diverse speaker and speech traits,â arXiv preprint arXiv:2505.14648, 2025. [25] T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, âUtmos: Utokyo-sarulab system for voicemos challenge 2022,â arXiv preprint arXiv:2204.02152, 2022. [26] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, âRobust speech recognition via large-scale weak supervision,â in International conference on machine learning. PMLR, 2023, p. 28 492â28 518. [27] C. Graham and N. Roll, âEvaluating openaiâs whisper asr: Performance analysis across diverse accents and speaker traits,â JASA Express Letters, vol. 4, no. 2, p. 025206, 02 2024. [Online]. Available: https://doi.org/10.1121/10.0024876