Paper deep dive
Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models
Pranav Tushar, Xiao Xiao Miao, Rong Tong
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 7/5/2026, 3:02:19 AM
Summary
This paper investigates child-centric voice anonymization by adapting a self-supervised learning (SSL) based pipeline to the child speech domain. The authors address the limitations of adult-centric systems by fine-tuning the HuBERT content encoder and HiFi-GAN vocoder using the MyST child speech corpus. The system is evaluated in both single-speaker and multi-speaker (including child-child and child-adult mixtures) scenarios. Results demonstrate that domain adaptation improves intelligibility and preserves perceived 'childness' while maintaining strong privacy protection. The study also highlights that in multi-speaker settings, utility is primarily constrained by the quality of target speaker extraction, especially in child-child mixtures.
Entities (10)
Relation Signals (4)
Pranav Tushar â affiliatedwith â Singapore Institute of Technology
confidence 100% ¡ Pranav Tushar ID 1 ... Singapore Institute of Technology
HuBERT â usedin â Voice Anonymization
confidence 100% ¡ via a HuBERT-based encoder [11]
HiFi-GAN â usedin â Voice Anonymization
confidence 100% ¡ reconstructs anonymized speech by feeding... into a HiFi-GAN vocoder
MyST Corpus â usedforadaptationof â HuBERT
confidence 90% ¡ fine-tune it on the MyST child speech corpus
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Voice anonymization aims to protect speaker identity while preserving linguistic content and speech usability. However, most anonymization systems are developed on adult speech, leading to degraded performance when applied to child speech. This paper investigates child-centric anonymization by adapting a self-supervised learning (SSL) based anonymization pipeline to the child speech domain. The system is adapted using child speech from the MyST corpus and evaluated under both single-speaker and two-speaker mixture conditions. Experimental results show that child-domain adaptation improves intelligibility and perceptual quality while maintaining strong privacy protection. Extending the approach to multi-speaker further demonstrates that combining target speaker extraction with child-adapted anonymization provides privacy protection while preserving conversational structure. These findings highlight the importance of child-specific adaptation for practical speech anonymization systems.
Tags
Links
- Source: https://arxiv.org/abs/2606.29897v1
- Canonical: https://arxiv.org/abs/2606.29897v1
Trouble viewing inline? Open PDF directly â
Full Text
33,391 characters extracted from source content.
Expand or collapse full text
Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models Pranav Tushar ID 1 , Xiao Xiao Miao ID 2,â , Rong Tong ID 1 1 Singapore Institute of Technology, Singapore 2 Duke Kunshan University, China tpranav2001@gmail.com, pranav.tushar@singaporetech.edu.sg, xiaoxiao.miao@dukekunshan.edu.cn, tong.rong@singaporetech.edu.sg Abstract Voice anonymization aims to protect speaker identity while pre- serving linguistic content and speech usability. However, most anonymization systems are developed on adult speech, leading to degraded performance when applied to child speech. This paper investigates child-centric anonymization by adapting a self-supervised learning (SSL) based anonymization pipeline to the child speech domain.The system is adapted using child speech from the MyST corpus and evaluated under both single-speaker and two-speaker mixture conditions. Experi- mental results show that child-domain adaptation improves in- telligibility and perceptual quality while maintaining strong pri- vacy protection. Extending the approach to multi-speaker fur- ther demonstrates that combining target speaker extraction with child-adapted anonymization provides privacy protection while preserving conversational structure. These findings highlight the importance of child-specific adaptation for practical speech anonymization systems. Index Terms: Childrenâs speech, voice privacy, anonymization, target speaker anonymization 1. Introduction Voice anonymization aims to suppress speaker identity while preserving linguistic content and downstream usability. The VoicePrivacy Challenge [1, 2, 3] has established common eval- uation protocols and reference systems, accelerating progress in both signal-processing and neural approaches [4, 5, 6, 7]. How- ever, most systems in this line of work are developed and evalu- ated on adult speech, and key components in modern pipelines, including content encoders and neural vocoders, are typically trained on adult data. Childrenâs speech constitutes a distinct acoustic domain, characterized by higher fundamental frequency, greater artic- ulatory and prosodic variability, and developmental disfluen- cies. These properties create a substantial mismatch for adult- trained anonymization pipelines. Recent evaluations show that while privacy protection can remain comparable under stan- dard automatic speaker verification (ASV) attackers, intelligi- bility and perceptual quality degrade substantially when adult- oriented systems are applied to child speech [8], motivating child-adapted anonymization. Additionally, existing approaches suffer from two limita- tions. First, many systems anonymize childrenâs speech by converting it toward adult-like voices, distorting age-dependent acoustic cues relevant to child-centered applications [8, 9]. Second, prior work considers only single-speaker scenarios, whereas real-world child speech, such as classroom interactions ** indicates the corresponding author. [10] or clinical sessions, frequently involves conversations be- tween children and adults. In such multi-speaker settings, it is desirable to selectively anonymize the target child while option- ally preserving or separately anonymizing the adult speaker. These limitations motivate child-to-child anonymization, where speaker identity is obfuscated while preserving per- ceived child-specific acoustic characteristics, producing natu- rally childlike anonymized speech under both single- and multi- speaker conditions. Specifically, we adopt a self-supervised learning (SSL)-based anonymization system [7] that (i) decom- poses the original speech into a content representation [11], a prosody representation using a pitch extractor, and a speaker representation [12]; (i) anonymizes the source speaker identity via a selective anonymization approach [13], where the source speaker embedding is replaced by a reference embedding drawn from a speaker pool; and (i) reconstructs anonymized speech by feeding the unchanged content and prosody representations together with the anonymized speaker embedding into a HiFi- GAN vocoder [14]. Since this system is originally trained on adult speech, we fine-tune both the content encoder and the vocoder on the MyST child speech corpus [15], and replace the adult speaker pool with a controlled child speaker pool con- structed from AI-generated voices screened for naturalness and age consistency, thereby improving utility on child speech and enabling child-to-child anonymization. We further extend this single-speaker system to multi-speaker conditions by combin- ing target speaker extraction with the anonymization pipeline. For single-speaker evaluation on both in-domain and zero-shot cross-accent benchmarks, the child-adapted system achieves strong privacy protection while maintaining compet- itive intelligibility and perceptual quality. For multi-speaker evaluation across three age-group pairings, adultâadult (A), childâadult (CA), and childâchild (C), anonymization remains effective across all overlap levels, intelligibility degradation is moderate for A mixtures but increases substantially for CA and C conditions, reflecting the additional challenge of child target speaker extraction. An overview of the proposed pipeline is shown in Fig. 1. The main contributions 1 of this work are: ⢠A systematic study of child-domain adaptation of SSL-based anonymization pipelines for child-to-child voice anonymiza- tion, evaluated on in-domain and zero-shot cross-accent benchmarks across privacy, intelligibility, perceptual quality, and perceived speaker type. ⢠Extension to two-speaker mixtures, demonstrating that pri- vacy remains relatively stable across overlap levels while in- telligibility is primarily constrained by target speaker extrac- 1 Source code, pretrained models, and audio samples are publicly available at https://github.com/pranavtushar/SSL-CVA. arXiv:2606.29897v1 [cs.SD] 29 Jun 2026 Figure 1: Child-centric anonymization. tion, particularly for childâchild mixtures. 2. Child-centric anonymization 2.1. SSL-based single-speaker child anonymization Given an input waveform x, we extract three disentangled rep- resentations: (i) soft content features c via a HuBERT-based encoder [11], (i) pitch contour f 0 , and (i) a speaker embed- ding s via ECAPA-TDNN [12]. The content and prosody rep- resentations are kept intact to preserve linguistic information, while the speaker embedding undergoes an identity transforma- tion to produce Ěs. A HiFi-GAN vocoder [14] then resynthe- sizes the anonymized waveform Ěx from (c, f 0 , Ěs). The base- line anonymization system is trained entirely on adult speech, leading to failure on child input at multiple stages. Prior work on child speech modeling reports consistent gains from domain adaptation of foundational speech models [16]. Motivated by this, we adapt the following components: HuBERT Content Encoder. The HuBERT soft content encoder, pretrained on adult speech, yields suboptimal linguistic repre- sentations for child utterances due to the acoustic mismatch in pitch range, formant frequencies, and speaking rate. We there- fore fine-tune it on the MyST child speech corpus [15] to im- prove the extraction of linguistically meaningful representations from child speech. Child Speaker Pool. In the original system, the source speaker embedding s is replaced by a reference embedding s ref drawn from an adult speaker pool, i.e. Ěs = s ref .For child-to- child anonymization, we replace this pool with a controlled set of child-like voices, where s ref is sampled to preserve age- consistent acoustic characteristics. The child speaker pool is constructed from AI-generated voices screened for naturalness and age consistency. Note that the ECAPA-TDNN speaker en- coder is kept fixed, as the anonymized identity is obtained by di- rectly replacing the source embedding with a reference embed- ding from the child speaker pool, which primarily determines the anonymized speaker identity. HiFi-GAN Vocoder. The HiFi-GAN vocoder, also trained on adult waveforms, tends to shift resynthesized child speech to- ward adult-like acoustics, distorting age-dependent cues such as pitch range and vocal tract resonances. We fine-tune it on the same MyST corpus to enable faithful resynthesis that preserves child-specific spectral and prosodic characteristics. 2.2. Target child anonymization in multi-speaker speech We extend child-centric anonymization to two-speaker mix- tures. Given a mixture x mix containing a target speaker and a non-target speaker, along with a short reference utterance Table 1: Datasets used for single speaker experiments. DatasetLanguageDomainUse MyST [15]English (US)In-domain Train+Dev: adapt Test: eval MPS [19]English (Indian acc. )Cross-accentZero-shot eval SpeechOcean [20]English (Mandarin acc.)Cross-accentZero-shot eval Table 2: Evaluation metrics and protocols for single and multi- speaker experiments. MetricBackendProtocol Single-speaker evaluation EER (â)ECAPA-TDNN ASV 2 [21]Speaker verification (original vs. anonymized) WER (â)Whisper Large-v3 3 [22]ASR error on anonymized speech NISQA-MOS (â)NISQA 4 [23]Objective speech quality Human eval.Listening testNaturalness, fluency, similarity, age Multi-speaker evaluation EER (â)ECAPA-TDNN ASV [21]Target verification (O: origâorig, OA: origâanon) tWER (â) gpt-4o-transcribe-diarizeTarget WER on diarized target channel DER (â)DiariZen + pyannote [24, 25]Diarization error vs. oracle RTTM r target , the goal is to anonymize only the target speakerâs iden- tity while leaving the non-target speech unmodified. The pipeline operates in three stages. First, a Conformer- based target speaker extraction (TSE) model [17, 18] isolates the target signal. The model estimates the target speakerâs complex spectrum from the STFT of x mix , conditioned on a speaker embedding derived from r target , and inverts it to ob- tain Ëx target . The non-target signal is recovered as the residual Ëx non-target = x mix â Ëx target . Second, the extracted target sig- nal Ëx target is anonymized using the single-speaker pipeline de- scribed in Section 2.1 with the child-adapted configuration for child targets and the base configuration for adult targets, pro- ducing Ěx target . Third, the anonymized mixture is reconstructed by replacing the target component Ěx mix = Ěx target +Ëx non-target . We evaluate across three age pairings: adultâadult (A), childâadult (CA), and childâchild (C) to reflect diverse real- world scenarios. This distinction is important because both extraction and anonymization become more challenging when speakers are acoustically similar, as in C pairs. 3. Experiments 3.1. Datasets and evaluation metrics Single-Speaker Datasets: Table 1 summarizes the datasets used for adaptation and evaluation. MyST [15] serves as the in- domain child speech corpus; we use the training partition for child-domain adaptation and reserve the test partition for evalu- ation. Zero-shot generalization is assessed on cross-accent En- glish datasets (MPS [19], SpeechOcean [20]). Multi-Speaker Mixtures: We construct two-speaker mixtures following a SparseLibriMix-style procedure [26]. Child speech is drawn from the MyST test split and adult speech from Lib- riSpeech test-clean. Utterances are cropped or padded to 5 s and mixed at 0 dB SNR with overlap ratios of 0â100% in 20% in- crements. The first speaker is designated as the anonymization target. We generate 50 mixtures per overlap level per age-group pairing (A,CA,C), yielding 900 mixtures in total. Evaluation Metrics: Table 2 summarizes the evaluation met- rics and protocols used in this work. For single-speaker ex- periments, we additionally conduct a listening study with 13 participants who rate naturalness, fluency, speaker similarity, and perceived age. Six source utterances are sampled across the evaluation datasets (two per dataset) to cover both short and longer child speech segments. For each source utterance, lis- Table 3: MyST (in-domain) component study for the selective anonymization system. Soft Encoder HiFi-GAN EER (â) WER (â) NISQA-MOS (â) BaseBase43.8017.313.60 FTBase38.1019.533.70 BaseFT40.6820.673.10 FTFT45.0916.643.36 teners hear the original signal and five anonymized variants and rate naturalness, fluency, speaker similarity, and perceived age. Each participant evaluates 30 anonymized samples in total. Formulti-speakerexperiments,pseudo- referencetranscriptsaregeneratedusing gpt-4o-transcribe-diarize 5 toobtainspeaker- attributed hypotheses for the target channel. Pseudo-references are used because ground-truth transcripts are not available after mixture construction, target extraction, and reconstruction. While this introduces a potential source of ASR error and may affect absolute WER values, it provides a consistent proxy for comparing relative intelligibility trends across systems and overlap conditions. 3.2. Experimental setup For child-domain adaptation, we fine-tune components of the SSL-based anonymization pipeline, namely the HuBERT con- tent encoder and the HiFi-GAN vocoder, on the MyST corpus following the strategy in [7]. We evaluate four configurations to analyze the effect of adaptation: Base/Base, FT/Base, Base/FT, and FT/FT. The reference speaker pool is constructed from AI- generated child-like voices using Typecast 6 and SpeechGen 7 , filtered for naturalness and age consistency. The pool con- tains 44 utterances from 16 synthetic child-like speakers. Sam- ples are manually screened to ensure naturalness and consistent child-like vocal characteristics before embedding extraction. In addition to the SSL-based anonymization approach, we also consider the signal-processing baseline B2 [4], which applies the McAdams coefficient and has been shown to better preserve child-specific characteristics [9]. The multi-speaker pipeline chains target speaker extraction, single-speaker anonymization, and mixture reconstruction. We employ a Conformer-based TSE model [17, 18] trained on adult speech without child adap- tation, allowing us to analyze domain mismatch effects under CA and C conditions. For child targets, we apply the child- adapted (FT/FT) anonymization configuration; for adult targets, the base configuration. 3.3. Results and discussion 3.3.1. Single-speaker anonymization. In-domain (MyST). Table 3 shows that adapting only one com- ponent (FT/Base or Base/FT) degrades performance, likely due to representation-synthesis mismatch. In contrast, full child- domain adaptation (FT/FT) achieves the best trade-off, improv- ing EER from 43.80% to 45.09% and reducing WER from 17.31% to 16.64%, although its NISQA-MOS is slightly lower than Base/Base. Therefore, we adopt the fully adapted (FT/FT) 5 https://developers.openai.com/api/docs/ models/gpt-4o-transcribe-diarize 6 https://typecast.ai/ 7 https://w.speechgen.app/ 0 2 4 Score MySTSpeechOcean B2 SSL-B SSL-FT 0 2 4 Score MPS Metric Naturalness Fluency Speaker_Similarity Age (child=1) Figure 2: Subjective evaluation across datasets: listener ratings for naturalness, fluency, speaker similarity & perceived child- ness. configuration for all subsequent evaluations, denoted as SSL- FT. Cross-domain generalization.Table 4 shows that SSL-FT achieves the strongest privacy protection, obtaining the high- est EER across all datasets. In terms of intelligibility, SSL- FT yields the lowest WER on MyST and MPS, and remains competitive on SpeechOcean, demonstrating that child-domain adaptation effectively preserves speech utility beyond the train- ing domain. Although SSL-FT does not consistently improve NISQA-MOS relative to SSL-B, it achieves stronger privacy protection and competitive intelligibility across datasets. Human evaluation. Fig. 2 summarizes subjective ratings across naturalness, fluency, and speaker similarity on a 1â5 Likert scale, where higher scores indicate better naturalness/fluency and greater perceived speaker similarity. For age perception, lis- teners performed a binary judgment (child vs. adult), and the re- ported value corresponds to the proportion of samples perceived as child speech (child = 1). The SSL-based anonymization sys- tems (SSL-B and SSL-FT) receive higher naturalness and flu- ency ratings than the B2 baseline, indicating improved percep- tual quality of the synthesized speech. Listener judgments fur- ther show that the child-adapted configuration (SSL-FT) more consistently preserves perceived childness while maintaining low speaker similarity scores.These results are consistent with the objective results, confirming that child-adapted neural anonymization can achieve effective identity concealment with- out substantially altering age-related acoustic characteristics. 3.3.2. Target children anonymization in multi-speaker speech Fig. 3 summarizes target-speaker privacy (EER), intelligibil- ity (tWER), and conversational-structure preservation (DER) across overlap ratios and age-group pairings. Privacy is robust to overlap.The first cluster of Fig. 3 shows the EER under different overlap ratios. After SSL-based anonymization, OA EERs (solid bars) are consistently and com- fortably higher than O EERs (hatched bars) across all age- group pairings, with EER increasing relative to the O condi- tion across all overlap ratios and subsets (A/CA/C), indicat- ing effective identity suppression even when the target speaker is first extracted from a mixture. Importantly, OA EER remains within a relatively narrow range as the overlap ratio increases from 0â100%, suggesting that anonymization effectiveness re- mains relatively stable across overlap levels. However, because anonymization operates on extracted target speech, privacy out- comes remain coupled to extraction quality. Table 4: Privacy (EER), intelligibility (WER), and perceptual quality (NISQA-MOS) across child speech datasets. Best results per dataset and metric are shown in bold. DatasetAge EER (OA) %âWER %âNISQA-MOSâ OrgB2SSL-B SSL-FTOrgB2SSL-B SSL-FTOrg B2 SSL-B SSL-FT MyST (test)age (8-11)15.3942.1043.8045.0914.5520.0217.3116.642.652.253.603.36 SpeechOcean age (6-10)4.3235.7134.9739.8827.2655.57 41.9943.363.372.293.312.85 age (11-15)2.2435.4237.7438.4612.0336.14 21.3823.613.402.313.382.97 MPSage (7-11)0.0131.4439.5040.9412.6818.3116.2015.722.151.753.312.65 0 10 20 30 40 1.3 5.5 8.0 22.4 18.8 8.6 0.8 4.9 8.1 27.8 20.5 9.8 0.0 4.7 7.0 28.1 14.0 7.1 1.0 2.3 5.9 28.5 16.3 8.0 2.0 3.4 8.1 30.7 20.5 9.4 2.2 2.7 7.1 25.4 21.7 9.1 A 0 20 40 60 22.4 9.7 14.8 30.3 42.8 19.0 16.5 11.6 14.6 30.9 43.2 18.4 20.2 15.6 15.7 35.2 53.7 17.9 19.8 10.8 16.0 30.6 54.5 17.6 12.3 12.4 13.6 32.5 47.2 12.7 15.8 8.4 18.4 33.5 38.7 17.8 CA EER tWER DER 0 20 40 60 80 19.2 14.4 30.5 28.3 62.2 35.2 23.5 13.5 26.9 34.4 66.0 36.2 24.1 12.4 25.0 32.6 53.5 32.6 21.9 6.5 24.2 32.5 72.1 34.4 22.4 11.3 23.1 31.1 81.5 32.6 18.6 15.3 20.9 31.5 64.9 30.8 C Overlap / Method 0% 20% 40% 60% 80% 100% Orig/O Anon/O Figure 3: Multi-speaker evaluation: tWER, DER, and EER (%) by age-group pairing (A/CA/C) and overlap ratio. O denotes originalâoriginal verification trials and OA denotes originalâanonymized verification trials. Utility degrades with separation difficulty, especially for childâ child mixtures. The second cluster of Fig. 3 shows that tWER increases with overlap and shows a strong dependence on age- group pairing. A mixtures remain the easiest regime (low baseline tWER and moderate increase after anonymization), while CA exhibits intermediate degradation. C mixtures are consistently the most challenging: even at low overlap, tWER is higher than A and CA, and it increases sharply under heavier overlap. DER follows a similar trend, with the largest diariza- tion errors observed in C conditions, reflecting the difficulty of separating acoustically similar child voices. The gap be- tween CA and C further emphasizes that child-specific acous- tics (higher pitch and rapid spectral dynamics) amplify extrac- tion ambiguity, making C the critical stress-test condition. Overall, the results highlight a privacyâutility decoupling in multi-speaker child scenarios: privacy metrics remain rela- tively stable across overlap levels, while downstream utility is primarily bounded by extraction quality, particularly for childâ child mixtures. This suggests that improving child-robust target speaker extraction is likely the most direct path to improving conversational child anonymization without weakening privacy. 4. Challenges and limitations Childrenâs speech differs substantially from adult speech, in- troducing challenges for both modeling and evaluation. Child recordings often contain higher background noise, classroom reverberation, microphone variability, and spontaneous vocal behaviors such as disfluencies and non-lexical vocalizations. These factors can create mismatches between transcripts and spoken content. As a result, ASR-based metrics such as WER may partially reflect transcription variability rather than purely anonymization-induced degradation. Another limitation concerns the adult-centric nature of sev- eral evaluation models used in the pipeline.Although the anonymization system is partially adapted to child speech, mul- tiple evaluation components remain primarily trained on adult data. These include ASR systems, diarization models, speaker verification attackers, and perceptual quality predictors. Conse- quently, some reported performance differences may reflect do- main mismatch in the evaluation models rather than limitations of the anonymization framework itself. For example, MOS predictors such as NISQA are largely trained on adult speech, which may reduce reliability when evaluating child voices. Age preservation is also difficult to quantify automatically. The age prediction models we tested generalized poorly on our child datasets, so we relied on human perceptual judgments to as- sess perceived age consistency; this improves validity but limits scalability. Finally, several components in the multi-speaker pipeline remain adult-trained, including the target speaker extraction model and the ASV attacker. This mismatch likely contributes to performance degradation in childâchild mixtures and under high overlap conditions. Stronger attacker models could further probe anonymization robustness. Overall, these factors high- light the need for child-adapted evaluation models and stronger attacker settings to better characterize privacyâutility trade-offs in child speech anonymization. 5. Conclusions and future work This work presented a child-centric study of voice anonymiza- tion through domain adaptation of self-supervised speech mod- els. Adapting the HuBERT content encoder and HiFi-GAN vocoder to child speech improves intelligibility while maintain- ing strong privacy protection across datasets. Human evalu- ations further indicate that the proposed approach better pre- serves perceived childness compared to signal-processing base- lines. Extending the analysis to multi-speaker mixtures reveals a clear decoupling between privacy and intelligibility. Privacy remains relatively stable across overlap conditions, while intel- ligibility degradation is primarily driven by target speaker ex- traction errors, particularly in childâchild mixtures. These findings highlight the importance of child-aware anonymization strategies as speech technologies are increas- ingly deployed in child-centered applications. Because chil- drenâs speech contains sensitive biometric information, privacy- preserving processing should be carefully considered when de- veloping speech technologies involving minors. Future work will explore child-adapted extraction and verification models, improved evaluation tools for child speech, and multilingual child speech anonymization using broader child speech corpora. 6. Acknowledgments This work is supported by the Singapore Ministry of Education (MOE) Academic Research Fund (AcRF) Tier 1 grant R-R13- A405-0005. 7. Generative AI Use Disclosure Generative AI tools were used in a limited manner for (i) lan- guage editing during manuscript preparation and (i) transcrip- tion support for WER-based evaluation using OpenAI ASR models. In addition, AI-generated text-to-speech (TTS) voices (Typecast and SpeechGen) were used to construct the child-like reference speaker pool for selective anonymization; these sam- ples were used only for reference embedding extraction and were not used to train the anonymization models. No gener- ative AI system was used for experimental design, model de- velopment, or interpretation of results. 8. References [1] N. Tomashenko, B. M. L. Srivastava, X. Wang, E. Vincent, A. Nautsch, J. Yamagishi, N. Evans, J. Patino, J.-F. Bonastre, P.-G. No Ě e, and M. Todisco, âIntroducing the VoicePrivacy Initiative,â in Proc. Interspeech 2020, 2020, p. 1693â1697. [2] M. Panariello, N. Tomashenko, X. Wang, X. Miao, P. Champion, H. Nourtel, M. Todisco, N. Evans, E. Vincent, and J. Yamag- ishi, âThe voiceprivacy 2022 challenge: Progress and perspec- tives in voice anonymisation,â IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, p. 3477â3491, 2024. [3] N. Tomashenko, X. Miao, P. Champion, S. Meyer, M. Panariello, X. Wang, N. Evans, E. Vincent, J. Yamagishi, and M. Todisco, âThe third voiceprivacy challenge: Preserving emotional expres- siveness and linguistic content in voice anonymization,â Com- puter Speech & Language, vol. 100, p. 101988, 2026. [4] J. Patino, N. Tomashenko, M. Todisco, A. Nautsch, and N. Evans, âSpeaker Anonymisation Using the McAdams Coefficient,â in Proc. Interspeech 2021, 2021, p. 1099â1103. [5] M. Panariello, F. Nespoli, M. Todisco, and N. Evans, âSpeaker anonymization using neural audio codec language models,â in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, p. 4725â 4729. [6] F. Fang, X. Wang, J. Yamagishi, I. Echizen, M. Todisco, N. Evans, and J.-F. Bonastre, âSpeaker anonymization using x-vector and neural waveform models,â in 10th ISCA Workshop on Speech Syn- thesis (SSW 10). ISCA, 2019, p. 155â160. [7] X. Miao, X. Wang, E. Cooper, J. Yamagishi, and N. Tomashenko, âLanguage-independent speaker anonymization approach using self-supervised pre-trained models,â in The Speaker and Lan- guage Recognition Workshop (Odyssey 2022).ISCA, 2022, p. 279â286. [8] A. Kulkarni, F. Teixeira, E. Hermann, T. Rolland, I. Trancoso, and M. M. Doss, âChildrenâs voice privacy: First steps and emerging challenges,â in Proc. Interspeech 2025, 2025, p. 2810â2814. [9] S. Dhar, S. R. Chetupalli, and P. Rao, âSpeaker anonymization for childrenâs oral reading assessment,â 2026. [10] P. Tushar, B. Zhang, I. Atmosukarto, D. Soh, R. Tong, and I. McLoughlin, âPersonalized ai-directed tutoring for oral pro- ficiency enhancement in language education,â Applied Sciences, vol. 16, no. 5, p. 2379, 2026. [11] B. Van Niekerk, M.-A. Carbonneau, J. Za Ě Äądi, M. Baas, H. Seut Ě e, and H. Kamper, âA comparison of discrete and soft speech units for improved voice conversion,â in ICASSP 2022-2022 IEEE In- ternational Conference on Acoustics, Speech and Signal Process- ing (ICASSP). IEEE, 2022, p. 6562â6566. [12] B. Desplanques, J. Thienpondt, and K. Demuynck, âEcapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,â in Proc. Interspeech 2020, 2020, p. 3830â3834. [13] B. M. L. Srivastava, M. Maouche, M. Sahidullah, E. Vincent, A. Bellet, M. Tommasi, N. Tomashenko, X. Wang, and J. Yam- agishi, âPrivacy and utility of x-vector based speaker anonymiza- tion,â IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, p. 2383â2395, 2022. [14] J. Kong, J. Kim, and J. Bae, âHifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,â Advances in neural information processing systems, vol. 33, p. 17 022â 17 033, 2020. [15] S. Pradhan, R. Cole, and W. Ward, âMy science tutor (myst)âa large corpus of childrenâs conversational speech,â in Proceedings of the 2024 Joint International Conference on Computational Lin- guistics, Language Resources and Evaluation (LREC-COLING 2024), 2024, p. 12 040â12 045. [16] R. Fan, N. Balaji Shankar, and A. Alwan, âBenchmarking chil- drenâs asr with supervised and self-supervised speech foundation models,â in Proc. Interspeech 2024, 2024, p. 5173â5177. [17] Y. Liu, X. Liu, X. Miao, and J. Yamagishi, âTarget speaker extrac- tion with curriculum learning,â in Proc. Interspeech 2024, 2024, p. 4348â4352. [18] â, âLibri2vox dataset: Target speaker extraction with di- verse speaker conditions and synthetic data,â arXiv preprint arXiv:2412.12512, 2024. [19] R. Gothi, R. Kumar, M. Pereira, N. Nayak, and P. Rao, âA dataset and two-pass system for reading miscue detection,â in Proc. In- terspeech 2024, 2024, p. 4014â4018. [20] J. Zhang, Z. Zhang, Y. Wang, Z. Yan, Q. Song, Y. Huang, K. Li, D. Povey, and Y. Wang, âspeechocean762: An open-source non- native english speech corpus for pronunciation assessment,â in Proc. Interspeech 2021, 2021. [21] J. S. Chung, A. Nagrani, and A. Zisserman, âVoxCeleb2: Deep Speaker Recognition,â in Proc. Interspeech 2018, 2018, p. 1086â 1090. [22] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, âRobust speech recognition via large-scale weak supervision,â in International conference on machine learning. PMLR, 2023, p. 28 492â28 518. [23] G. Mittag, B. Naderi, A. Chehadi, and S. M Ě oller, âNISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets,â in Proc. Inter- speech 2021, 2021, p. 2127â2131. [24] J. Han, F. Landini, J. Rohdin, A. Silnova, M. Diez, J. Ë Cernock ` y, and L. Burget, âFine-tune before structured pruning: Towards compact and accurate self-supervised models for speaker diariza- tion,â in Proc. Interspeech 2025, 2025, p. 1583â1587. [25] H. Bredin, âpyannote.audio 2.1 speaker diarization pipeline: prin- ciple, benchmark, and recipe,â in Proc. Interspeech 2023, 2023, p. 1983â1987. [26] J. Cosentino, M. Pariente, S. Cornell, A. Deleforge, and E. Vin- cent, âLibrimix: An open-source dataset for generalizable speech separation,â arXiv preprint arXiv:2005.11262, 2020. [27] M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Son- deregger, âMontreal forced aligner: Trainable text-speech align- ment using kaldi.â in Interspeech, vol. 2017, 2017, p. 498â502. A. Additional Experimental Details A.1. Comparison of WER The multi-speaker experiments in the main paper evaluate target-speaker intelligibility using pseudo-reference transcripts generated by gpt-4o-transcribe-diarize. To verify that the reported conclusions are not dependent on the choice of reference transcripts, we additionally compute WER using transcript-grounded references derived from Montreal Forced Aligner (MFA) [27]. A.1.1. Protocol We generated 900 deterministic two-speaker mixtures across adultâadult (A), childâadult (CA), and childâchild (C) con- ditions at overlap ratios of 0%, 20%, 40%, 60%, 80%, and 100%.Whisper Large-v3 was used to transcribe the anonymized target channel (target anon.wav) for both evaluation protocols. The only difference lies in the refer- ence transcript: (i) Pseudo Ref., which uses pseudo-reference transcripts generated by gpt-4o-transcribe-diarize (as reported in the main paper), and (i) GT Ref., which uses transcript-grounded references derived from the original source utterances using MFA. Prior to WER computation, transcripts were normalized by removing punctuation and annotation arti- facts, expanding common contractions, and standardizing nu- meric expressions where applicable. Because Whisper occa- sionally produced long hallucinated hypotheses for short child- speech segments, per-utterance WER was capped at 100% be- fore averaging within each subset and overlap condition. A.1.2. Discussion Table 5 shows that although the absolute WER values differ between the pseudo-reference and MFA transcript-grounded evaluation protocols, both exhibit the same qualitative trends. Across all overlap conditions, A consistently achieves the lowest WER, whereas CA and C remain substantially more challenging. In particular, C yields the highest WER across most overlap settings, indicating that target speaker extraction becomes increasingly difficult when both speakers share similar child-specific acoustic characteristics. The consistency between the two evaluation protocols provides additional evidence that the conclusions reported in the main paper are not driven by the pseudo-reference generation procedure. Table 5: Comparison of target-speaker WER (%) computed using pseudo-reference (Pseudo Ref.) and MFA transcript- grounded (GT Ref.) reference transcripts. The pseudo-reference WER values are reproduced from Fig. 3 of the main paper for ease of comparison. Although the absolute values differ, both evaluation protocols exhibit similar qualitative trends across A, CA, and C conditions. Lower values indicate better in- telligibility (â). AACACC Overlap (%) Pseudo Ref. GT Ref.Pseudo Ref. GT Ref.Pseudo Ref. GT Ref. 018.820.8542.841.3562.259.15 20 20.520.5943.248.7666.064.28 40 14.017.7653.749.2653.559.91 6016.320.0954.540.9872.169.35 8020.522.7447.243.4281.571.38 100 21.723.6138.741.5764.966.54