Paper deep dive
Acoustic and perceptual differences between standard and accented Chinese speech and their voice clones
Tianle Yang, Chengzhe Sun, Phil Rose, Siwei Lyu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/2/2026, 11:52:51 PM
Summary
This paper investigates the impact of accent variation on voice cloning performance by comparing standard and heavily accented Mandarin speech. Using a combination of computational analysis (ECAPA-TDNN speaker embeddings) and human perception experiments, the authors find that while embedding-based metrics show no reliable difference in original-clone distances between accent types, human listeners perceive lower similarity for accented clones. Conversely, voice cloning provides a greater intelligibility gain for accented speech compared to standard speech, suggesting that accent attenuation occurs during the cloning process.
Entities (7)
Relation Signals (3)
Mandarin Heavy Accent Speech Corpus â sourceof â Accented Mandarin speech
confidence 95% ¡ The accented Mandarin speech was drawn from the Mandarin Heavy Accent Speech Corpus
ECAPA-TDNN â usedfor â Speaker Embedding Extraction
confidence 95% ¡ We represent each speech token using deep speaker embeddings from an ECAPA-TDNN model
ElevenLabs â evaluatedin â Voice Cloning
confidence 90% ¡ We generated cloned speech using three commercial voice-cloning systems: ElevenLabs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Voice cloning is often evaluated in terms of overall quality, but less is known about accent preservation and its perceptual consequences. We compare standard and heavily accented Mandarin speech and their voice clones using a combined computational and perceptual design. Embedding-based analyses show no reliable accented-standard difference in original-clone distances across systems. In the perception study, clones are rated as more similar to their originals for standard than for accented speakers, and intelligibility increases from original to clone, with a larger gain for accented speech. These results show that accent variation can shape perceived identity match and intelligibility in voice cloning even when it is not reflected in an off-the-shelf speaker-embedding distance, and they motivate evaluating speaker identity preservation and accent preservation as separable dimensions.
Tags
Links
- Source: https://arxiv.org/abs/2604.01562v1
- Canonical: https://arxiv.org/abs/2604.01562v1
Trouble viewing inline? Open PDF directly â
Full Text
29,036 characters extracted from source content.
Expand or collapse full text
Acoustic and perceptual differences between standard and accented Chinese speech and their voice clones Tianle Yang 1,â , Chengzhe Sun 2 , Phil Rose 3 , Siwei Lyu 2 1 Department of Linguistics, University at Buffalo, United States 2 Department of Computer Science and Engineering, University at Buffalo, United States 3 Emeritus Faculty, Australian National University, Australia tianleya@buffalo.edu Abstract Voice cloning is often evaluated in terms of overall quality, but less is known about accent preservation and its perceptual con- sequences. We compare standard and heavily accented Man- darin speech and their voice clones using a combined computa- tional and perceptual design. Embedding-based analyses show no reliable accented-standard difference in original-clone dis- tances across systems. In the perception study, clones are rated as more similar to their originals for standard than for accented speakers, and intelligibility increases from original to clone, with a larger gain for accented speech. These results show that accent variation can shape perceived identity match and intelli- gibility in voice cloning even when it is not reflected in an off- the-shelf speaker-embedding distance, and they motivate eval- uating speaker identity preservation and accent preservation as separable dimensions. Index Terms: voice cloning, audio deepfakes, accent variation, speech intelligibility, speaker embeddings 1. Introduction Accents 1 are a natural consequence of dialectal variation. When speech is evaluated against a socially defined standard, system- atic phonetic differences and the phonological patterns that or- ganize them are often labeled as âaccentâ [1]. In everyday com- munication, accents can affect intelligibility, social judgments, and access to services. In modern speech technology, they also raise a basic scientific question: when a system learns a voice, what parts of accent-related structure does it preserve, and what parts does it reduce or reshape? This question becomes especially important in voice cloning (audio deepfakes). Voice cloning has clear benefits for speech interfaces, dubbing, accessibility tools, and personalized speech avatars. At the same time, its increasing realism raises security and ethical concerns because cloned voices can be used for impersonation [2], fraud [3], and deception [4]. Much exist- ing work therefore emphasizes detection and robustness [5, 6], speaker modeling and discrimination [7, 8, 9], or artifacts intro- duced by generative models [10, 11, 12, 13]. By contrast, we still have limited understanding of how accent variation itself is represented in voice-cloned speech, and how cloned accented speech affects human perception. This gap matters both for safety (how âauthenticâ a clone is with respect to accent) and for potential applications (whether cloning can improve intelli- gibility for heavily accented speech). ** indicates the corresponding author. 1 Linguistically, all varieties have accents; here accented is used in the common sense (regional features relative to standard Mandarin). A useful way to approach this problem is to ask how accent- related changes appear in speaker representations learned by deep models. Speaker embedding systems such as ECAPA- TDNN [14] are designed to capture speaker-discriminative in- formation, and they are widely used in speaker verification and related tasks [15]. If voice cloning modifies accent-specific cues in a consistent way, then such modifications may shift how orig- inal and cloned utterances relate in embedding space. However, it is not clear how large these shifts are, whether they differ for accented versus standard speech, or whether they depend on the voice-cloning system. Quantifying these patterns provides a concrete measurement of how âspeaker-likeâ a clone is relative to its source speech under different accent conditions. In this paper, we investigate speaker identity preservation and potential accent attenuation in modern commercial voice cloning by combining computational analysis with a controlled perception experiment. We focus on Mandarin speech produced with a heavy regional accent and a socially defined standard ac- cent. Our goal is not to benchmark overall naturalness or syn- thesis quality, but to characterize how accent-related variation is preserved or reshaped by voice cloning, and how any such changes relate to human perception. Specifically, we address the following questions: 1. Accent and embedding divergence: Do voice clones de- rived from heavily accented speakers diverge from their source speech in embedding-based speaker space more (or less) than clones derived from standard-accented speakers? 2. System dependence: Are any accent-related differences in embedding divergence consistent across voice-cloning sys- tems, or do different systems exhibit different degrees or di- rections of accent attenuation? 3. Perceptual consequences: If accent attenuation is observed, does voice cloning improve intelligibility for heavily ac- cented speech while maintaining perceived speaker identity? By answering these questions, we provide a unified evalu- ation of accent preservation in voice cloning that links model- internal speaker representations to perceptual outcomes. 2. Experimental setting In this study, we conduct two experiments: a computational analysis of embedding-based speaker distances, and a percep- tual study examining intelligibility and speaker similarity for accented and standard Mandarin speech and their voice clones. 2.1. Voice-clone generation The accented Mandarin speech was drawn from the Mandarin Heavy Accent Speech Corpus [16], while the standard Man- arXiv:2604.01562v1 [cs.SD] 2 Apr 2026 darin speech was drawn from AISHELL-3 [17]. We generated cloned speech using three commercial voice-cloning systems: ElevenLabs [18], MiniMax [19], and AnyVoice [20]. For each speaker and system, we used a fixed enrollment duration of ap- proximately 20 seconds to ensure comparability across systems; the same enrollment audio segment (after identical preprocess- ing) was used across all three systems for a given speaker. To reduce nuisance variability across recordings, all enroll- ment and source recordings were processed with a shared pre- processing pipeline prior to cloning and analysis. Specifically, audio was converted to mono, resampled to a common sampling rate (16 kHz), and amplitude-normalized; we did not apply ad- ditional denoising or manual editing beyond these standardiza- tion steps. For synthesis, we used a fixed set of text prompts shared across systems. All systems were run with fixed, com- parable synthesis settings (using default system settings where platform-specific controls were not directly comparable). All synthesized outputs were exported and then standard- ized again before downstream analysis, using the same mono conversion, resampling, and amplitude-normalization proce- dure described above. Importantly, the original and cloned stimuli used for the embedding-distance analysis were identi- cal to those used in the perception experiment. This design provides tight experimental control and ensures that any differ- ences between computational and perceptual outcomes cannot be attributed to differences in the audio materials. 2.2. Speaker Embedding Extraction and Cosine-Distance Computation We represent each speech token using deep speaker embed- dings from an ECAPA-TDNN model [14]. Specifically, we use the SpeechBrain toolkit [21] with the pretrained checkpoint re- leased on Hugging Face [22], which outputs a 192-dimensional embedding vector [23]. All tokens are loaded as waveforms using PyTorch and torchaudio [24, 25]. Following the check- point specification, inputs are converted to mono and resam- pled to 16 kHz [22]. Each token waveform is then passed to encode batch via SpeechBrainâs EncoderClassifier inference interface [26] to obtain a token-level embedding. We L2-normalize embeddings before distance computation to re- move magnitude effects, ensuring cosine distance primarily re- flects angular differences and is comparable across tokens. To obtain multiple embedding observations per speaker while keeping comparisons consistent across original and cloned speech, we segment each file into fixed-length tokens using a sliding window (token length L = 3.0 s; hop size H = 1.5 s). To reduce the influence of silence, we apply a sim- ple energy-based VAD computed on 30 ms frames with a 10 ms frame shift. A frame is labeled as speech if its RMS energy is greater than â35 dB relative to the file-level maximum RMS. We retain a token only if at least 60% of its frames are labeled as speech. To avoid domination by speakers with many usable tokens, we cap the number of retained tokens at 120 per speaker per condition (original vs. cloned) by random subsampling. For any two embeddingse i ande j , cosine similarity is cos(e i ,e j ) = e T i e j âĽe i âĽe j ⼠,(1) and cosine distance is defined as d cos (e i ,e j ) = 1â cos(e i ,e j ).(2) Within each speaker, we compute three distance sets from available token pairs: originalâoriginal distances (d O ) among original tokens, cloneâclone distances (d C ) among cloned to- kens, and originalâclone distances (d OC ) between original and cloned tokens. We summarize each set by its mean (e.g.,d OC ) as the primary speaker-level metric. To control computational cost, we use all token pairs when feasible; when the number of pairs is large, we randomly sample up to 25,000 within- condition pairs and up to 50,000 cross-condition pairs (originalâ clone) per speaker to estimate the means. Inferential statistics are conducted on speaker-level distances (one aggregated value per speaker per condition), avoiding pseudo-replication from treating token-pair distances as independent observations. Fi- nally, to quantify whether cloning increases within-speaker dis- persion in embedding space, we define a speaker-level clone- divergence measure as â div = d OC âd O ,(3) where positive values indicate that originalâclone distances ex- ceed the baseline originalâoriginal within-speaker distances. 2.3. Participants Participants were recruited for an online perception experiment administered in Qualtrics [27]. All participants reported being native speakers of Mandarin Chinese who use it as their pri- mary daily language. They also reported normal hearing and no history of speech or hearing disorders. The study protocol was approved by the Institutional Review Board (IRB), and all participants provided informed consent and participated volun- tarily. We collected basic demographic information (age group and gender) for descriptive purposes. A total of N = 138 participants completed the study. We excluded participants who did not complete the full experiment, reported listening without headphones, or were identified as du- plicate submissions. After exclusions, N = 67 participants were retained for analysis (age group: 18â38, n = 43; 39â59, n = 15; 60+, n = 9; gender: female, n = 40; male, n = 27). 3. Results 3.1. Embedding-based speaker distances for accented and standard speech and their voice clones Figure 1 visualizes model-level ECAPA-TDNN cosine dis- tances between original speech and the corresponding voice clones for standard and accented Mandarin speakers. The top panel shows the mean originalâclone distance ( d OC ) for each commercial system, with error bars reflecting 95% confidence intervals over speakers. The bottom panel reports the clone- divergence measure (â div =d OC âd O ), where the dashed horizontal line indicates the zero point (i.e., originalâclone dis- tances equal the within-original baseline). Visually, the clone- divergence estimates for AnyVoice and MiniMax lie close to the zero reference line. This figure is not intended as a performance benchmark across systems; instead, it illustrates that under our controlled tokenization and balancing procedure, some systemâ condition combinations yield only small deviations from the within-original baseline in embedding space. To further quantify the patterns suggested by the descriptive results above, we tested whether (i) originalâcloned distances reliably exceed the within-original baseline (clone divergence), and (i) originalâcloned distances differ between accented and standard speech within each voice-cloning system. All tests were performed on speaker-level mean cosine dis- tances to avoid pseudo-replication from treating token-pair dis- tances as independent observations. For each speaker in each 0.0 0.2 0.4 0.6 Originalâclone distance Standard Accented 0.0 0.1 0.2 0.3 ElevenLabsAnyVoiceMiniMax Clone divergence Standard Accented Figure 1: Model-level ECAPA-TDNN cosine distances be- tween original speech and voice clones for standard and ac- cented Mandarin speakers. Top: mean originalâcloned dis- tance ( d OC ). Bottom: clone divergence (â div =d OC âd O ); dashed line marks â div = 0. Error bars show 95% confidence intervals across speakers. dataset, we computed the mean cosine distance for originalâ original pairs (O) and for originalâcloned pairs (OC). We then fitted a linear mixed-effects model (REML) predicting mean distance from dataset (standard vs. accented), voice-cloning system (ElevenLabs, AnyVoice, MiniMax), and distance type (O vs. OC), including all interactions, with a random inter- cept for speaker within dataset. Degrees of freedom were ap- proximated using Satterthwaiteâs method. Because higher-order interactions are present, inferential conclusions are based on planned contrasts from estimated marginal means rather than on individual model coefficients. Holm correction was applied separately within each planned-contrast family. We first tested clone divergence (OCâO) within each dataset-by-system condition.Clone divergence was signifi- cant for ElevenLabs in both datasets (standard: estimate = 0.21118, SE = 0.03275, Holm-adjusted p = 8.82 Ă 10 â5 ; accented: estimate = 0.21193, SE = 0.03275, Holm-adjusted p = 8.82Ă 10 â5 ), indicating that originalâcloned distances ex- ceed the within-original baseline by a similar amount for stan- dard and accented speech. In contrast, AnyVoice and Min- iMax showed no reliable clone divergence in either dataset (AnyVoice, standard: estimate = 0.00892, SE = 0.03782, Holm-adjusted p = 1.000; AnyVoice, accented: estimate = 0.00382, SE = 0.03782, Holm-adjusted p = 1.000; Min- iMax, standard: estimate = 0.01725, SE = 0.03782, Holm- adjusted p = 1.000; MiniMax, accented: estimate = 0.00640, SE = 0.03782, Holm-adjusted p = 1.000). We next tested whether accented speech differs from stan- dard speech in originalâcloned distances (accentedâstandard, evaluated at OC) within each system. In all three systems, the accented condition showed numerically larger originalâcloned distances than the standard condition, but these differences were not statistically reliable after multiple-comparison correction (ElevenLabs: estimate = 0.09018, SE = 0.04743, Holm- adjusted p = .2113; AnyVoice: estimate = 0.06242, SE = 0.05476, Holm-adjusted p = .4080; MiniMax: estimate ElevenLabsAnyVoiceMiniMax standardaccentedstandardaccentedstandardaccented 1 2 3 4 5 Similarity rating (1â5) Speaker similarity (clone relative to original): distribution ElevenLabsAnyVoiceMiniMax standardaccentedstandardaccentedstandardaccented 1 2 3 4 5 Similarity rating (1â5) Speaker similarity (clone relative to original): paired by participant Figure 2: Listener-rated speaker similarity between each voice clone and its corresponding original recording for standard and accented Mandarin speakers, shown separately for different systems. Top: distributions of participant-level mean similarity ratings. Bottom: within-participant paired means for standard vs. accented. = 0.07170, SE = 0.05476, Holm-adjusted p = .4080). Overall, the results show clear system-dependent differ- ences in clone divergence, but no reliable evidence that accent systematically changes originalâcloned embedding distances in this dataset. 3.2. Perceptual evaluation of speaker similarity and intel- ligibility for standard and accented speech and their voice clones Figure 2 summarizes listenersâ similarity ratings between each voice clone and its corresponding original recording, separately for the standard and accented speaker sets and for each commer- cial system (ElevenLabs, AnyVoice, MiniMax). The top row shows the distribution of participant-level mean similarity rat- ings within each condition, while the bottom row shows within- listener changes across the two speaker sets 2 . Two qualitative patterns stand out. First, similarity rat- ings are generally higher for AnyVoice and MiniMax than for ElevenLabs, with many ratings near the upper end of the scale in the standard condition, aligning with our earlier acoustic analysis, which indicated smaller originalâclone divergences for AnyVoice and MiniMax. Second, across systems, per- ceived similarity between each speakerâs original recording and their corresponding voice clone appears lower for the accented speaker set than for the standard speaker set. Notably, this visual trend does not align with our earlier acoustic analysis, which did not show a reliable accent-related change in origi- nalâclone divergence; we therefore test it formally below. We analysed similarity ratings using a cumulative link mixed model (CLMM; logit link) for the 5-point ordinal scale. The fixed-effects structure included speaker set (standard vs. ac- cented), system (ElevenLabs, AnyVoice, MiniMax), and their interaction. The model included random intercepts for listener 2 Note that values on the plot may fall between integers because each point summarizes a participantâs mean rating across multiple stimuli within the condition. ElevenLabsAnyVoiceMiniMax standardaccentedstandardaccentedstandardaccented â2 0 2 4 Gain (clone â original) Intelligibility gain (clone â original): distribution ElevenLabsAnyVoiceMiniMax standardaccentedstandardaccentedstandardaccented â2 0 2 4 Gain (clone â original) Intelligibility gain (clone â original): paired by participant Figure 3: Listener-rated intelligibility gain for each clone relative to its matched original recording, shown separately for speaker sets and for each system.Top: distributions of participant-level mean gains. Bottom: within-participant paired mean gains for Standard vs. Accented. and stimulus; adding a by-listener random slope for speaker set did not improve fit (AIC RI = 3132.05 vs. AIC RS = 3134.65), so we report the random-intercept model. System effects were large: relative to ElevenLabs, similarity ratings were higher for AnyVoice ( Ë Î˛ = 3.34, z = 6.81, p < .001) and MiniMax ( Ë Î˛ = 3.51, z = 7.14, p < .001). Within each system, perceived similarity between each original recording and its corresponding clone was lower for the accented speaker set than for the standard set in AnyVoice (standardâaccented on the latent logit scale: â = 1.32, z = 2.54, Holm-adjusted p = .011) and MiniMax (â = 1.22, z = 2.35, Holm-adjusted p = .019), while the same con- trast was not reliable for ElevenLabs (â = 0.23, z = 0.53, p = .600). Across systems, the overall standardâaccented contrast was also significant when marginalizing over systems (equal-weight averaging: p = .0008; proportional-weight av- eraging: p = .0018). In a complementary omnibus likelihood- ratio test, adding the dataset-related terms (main effect plus in- teractions) significantly improved model fit (Ď 2 (3) = 10.23, p = .0167). Overall, listeners rated clones as more similar to originals for standard than accented speakers, and similarity was higher for AnyVoice and MiniMax than for ElevenLabs. For intelligibility, we summarize each listenerâs judgments using an intelligibility gain score that compares a clone di- rectly to its matched original. For each participant, system, and speaker set (standard vs. accented), we first average the 1â5 in- telligibility ratings across items separately for the original and the clone, and then compute gain as â intell = r clone âr orig . Using gain has two advantages: it provides a within-listener, within-item baseline correction that reduces individual differ- ences in rating scale use, and it targets the quantity of interest for voice cloning, namely how much intelligibility changes rel- ative to the original rather than absolute ratings. Figure 3 visualizes these gains by system and speaker set. The distributional summaries (top) and within-participant paired means (bottom) indicate that gains are generally positive and tend to be larger for the accented speaker set than for the standard set, with the strength of this pattern varying by system. We analysed intelligibility ratings using the same CLMM specification as for similarity. Relative to originals, clones received higher intelligibility ratings on the latent logit scale ( Ë Î˛ = 2.02, z = 8.48, p < 2Ă 10 â16 ). Intelligibility ratings were lower for the accented than for the standard speaker set ( Ë Î˛ = â2.12, z = â2.77, p = .0056). Importantly, the clone advantage was larger for accented than for standard speech, as indicated by a significant source-by-speaker-set interaction ( Ë Î˛ = 1.53, z = 4.79, p = 1.65Ă 10 â6 ). No system-related effects were reliable (all p⼠.128), suggesting that this pattern does not differ clearly across ElevenLabs, AnyVoice, and Min- iMax. Overall, clones were rated as more intelligible than their matched originals, and this cloneâoriginal advantage was larger for accented than for standard speech. 4. Discussion and conclusion This study links embedding-based speaker distances to human perception for standard and accented Mandarin speech and their voice clones across commercial systems. In the embedding analysis, we did not find evidence that originalâclone distances differ between the standard and accented speaker sets. By con- trast, perception showed a consistent asymmetry: similarity rat- ings were higher for standard originals relative to their matched clones than for accented originals relative to their matched clones, while intelligibility showed a larger clone advantage for accented speech. These outcomes are not contradictory because the com- putational and perceptual measures target different constructs. ECAPA-TDNN distances reflect similarity in a speaker- discriminative representation that may downweight accent- related cues. Thus, a null accented-standard effect in embed- ding distances does not imply that accent is unchanged by cloning; it implies that any accent-related modification is small or weakly represented in this embedding space. Listeners, how- ever, may treat accent cues as part of the evidence for speaker identity match, which can yield a standard-accented asymmetry in perceived similarity even when a generic speaker-embedding metric is insensitive to it. These similarity results suggest that clones of accented speech are perceived as less faithful to their source than clones of standard speech. The intelligibility results point in the opposite direction with respect to benefit: cloning increases intelligibility over- all, and the increase is larger for accented speech. Together, these patterns are consistent with an accent-attenuation account: cloning may shift accented speech toward more standard-like realizations that improve comprehension, but this shift reduces perceived similarity to the original when accent is salient for identity judgments. Practically, these findings motivate separating speaker preservation from accent preservation in evaluation. Future work could add accent-sensitive acoustic measures and explicit accent ratings for originals and clones, so that accent change can be linked to specific acoustic differences and their percep- tual consequences. Future work could also test a broader range of cloning models, more languages, and platform settings, since the degree of accent attenuation might vary with model choice, different configurations, or language-specific features. Overall, our results indicate that the linguistic identity of a speaker, such as accent variations, can affect perceived identity match and intelligibility in voice cloning even when it is not reflected in an off-the-shelf speaker-embedding distance. 5. References [1] J. K. Chambers and P. Trudgill, Dialectology.Cambridge Uni- versity Press, 1998. [2] B. Finley, âDeepfake of principalâs voice is the latest case of AI being used for harm,â AP News, 2024. [3] C. Stupp, âFraudsters used AI to mimic CEOâs voice in unusual cybercrime case,â The Wall Street Journal, vol. 30, no. 08, 2019. [4] S. Rao, A. K. Verma, and T. Bhatia, âA review on social spam detection: Challenges, open issues, and future directions,â Expert Systems with Applications, vol. 186, p. 115742, 2021. [5] C. Sun, T. Yang, and S. Lyu, âA survey of AI-generated voices and their detection,â APSIPA Transactions on Signal and Information Processing, 2026, accepted 3 December 2025. [6] B. Nguyen and T. Le, âAnalyzing Reasoning Shifts in Audio Deepfake Detection under Adversarial Attacks: The Reasoning Tax versus Shield Bifurcation,â arXiv preprint arXiv:2601.03615, 2026. [7] T. Yang, C. Sun, S. Lyu, and P. Rose, âForensic deepfake audio detection using segmental speech features,â Forensic Science International, vol. 379, p. 112768, 2026. [On- line]. Available: https://w.sciencedirect.com/science/article/ pii/S0379073825004128 [8] W. Chen and X. Jiang, âVoice-cloning artificial-intelligence speakers can also mimic human-specific vocal expression,â 2023. [9] W. Chen, M. D. Pell, and X. Jiang, âHuman and AI voice iden- tities evoke shared neural signatures during speaker recognition across changes in speech content and prosody,â bioRxiv, p. 2025â12, 2025. [10] E. Coletta, D. Salvi, V. Negroni, D. U. Leonzio, and P. Bestagini, âAnomaly detection and localization for speech deepfakes via fea- ture pyramid matching,â arXiv preprint arXiv:2503.18032, 2025. [11] D. Salvi, âData-driven techniques for speech and multimodal deepfake detection,â 2024. [12] X. Guo, Y. Xie, H. Cheng, J. Zhou, J. Liu, H. Huang, L. Ye, and Q. Zhang, âTowards Explicit Acoustic Evidence Perception in Audio LLMs for Speech Deepfake Detection,â arXiv preprint arXiv:2601.23066, 2026. [13] T. Yang, C. Sun, P. Rose, C. L. Jacobs, and S. Lyu, âAssessing the ability of neural tts systems to model consonant-induced f0 perturbation,â Computer Speech & Language, vol. 100, p. 101983, 2026. [14] B. Desplanques, J. Thienpondt, and K. Demuynck, âEcapa- tdnn:Emphasized channel attention, propagation and ag- gregation in tdnn based speaker verification,â arXiv preprint arXiv:2005.07143, 2020. [15] C.-L. Huang, âSpeaker Characterization Using TDNN, TDNN- LSTM, TDNN-LSTM-Attention based Speaker Embeddings for NIST SRE 2019.â in Odyssey, 2020, p. 423â427. [16] OpenDataLab,âMandarinHeavyAccentConversational SpeechCorpus,âhttps://opendatalab.com/OpenDataLab/ Mandarin HeavyAccentConversationaletc,n.d.,accessed: 2025-05-19. [17] Y. Shi, H. Bu, X. Xu, S. Zhang, and M. Li, âAISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines,â 2015. [Online]. Available: https://arxiv.org/abs/2010.11567 [18] Elevenlabs, âElevenLabs - Generative AI Text to Speech & Voice Cloning,â https://elevenlabs.io, 2025. [19] Minimax, âMinimax - Generative AI Text to Speech & Voice Cloning,â https://w.minimax.io/audio/voices-cloning, 2025. [20] Anyvoice, âAnyvoice - Generative AI Text to Speech & Voice Cloning,â https://anyvoice.net, 2025. [21] M. Ravanelli, T. Parcollet, P. Plantinga, A. Rouhe, S. Cornell, L. Lugosch, C. Subakan, N. Dawalatabad, A. Heba, J. Zhong, J.-C. Chou, S.-L. Yeh, S.-W. Fu, C.-F. Liao, E. Rastorgueva, F. Grondin, W. Aris, H. Na, Y. Gao, R. D. Mori, and Y. Bengio, âSpeechBrain: A General-Purpose Speech Toolkit,â 2021. [22] SpeechBrain,âspeechbrain/spkrec-ecapa-voxceleb:Speaker VerificationwithECAPA-TDNNembeddingsonVox- Celeb(modelcard),âhttps://huggingface.co/speechbrain/ spkrec-ecapa-voxceleb, 2021. [23] â, âspeechbrain/spkrec-ecapa-voxceleb: hyperparams.yaml (linneurons=192),âhttps://huggingface.co/speechbrain/ spkrec-ecapa-voxceleb/blob/main/hyperparams.yaml, 2024. [24] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K Ě opf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chil- amkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, âPy- Torch: An Imperative Style, High-Performance Deep Learning Library,â in Advances in Neural Information Processing Systems (NeurIPS), 2019, p. 8024â8035. [25] PyTorchTeam,âtorchaudiotutorial:AudioResam- pling,âhttps://docs.pytorch.org/audio/stable/tutorials/ audio resamplingtutorial.html, 2025. [26] SpeechBrain, âSpeechBrain Inference API: EncoderClassifier (encodebatch, classifybatch),â https://speechbrain.readthedocs. io/en/latest/API/speechbrain.inference.classifiers.html, 2024. [27] Qualtrics, âQualtrics XM: Online Survey Platform,â https://w. qualtrics.com.