Paper deep dive
Cross-Dataset, Age, and Gender Generalization: A Comprehensive Analysis of Fine-Tuning Strategies for Low-Resource Children's ASR
Paban Sapkota, Hemant Kumar Kathania, Mikko Kurimo, Sudarsana Reddy Kadiri, Shrikanth Narayanan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 6/21/2026, 4:35:12 AM
Summary
This paper investigates fine-tuning strategies for Self-Supervised Learning (SSL) models to improve Automatic Speech Recognition (ASR) for children's speech. The study evaluates three models (Wav2Vec2, HuBERT, and WavLM) across two datasets (PFSTAR and CMU Kids) focusing on age-specific, gender-specific, and cross-dataset generalization. Key findings include that fine-tuning on younger children improves generalization to older children, fine-tuning helps mitigate male-preference bias in pre-trained models, and cross-dataset performance suffers significantly due to accent and demographic mismatches.
Entities (8)
Relation Signals (6)
Wav2Vec2 â evaluatedon â PFSTAR
confidence 100% ¡ Wav2Vec2 achieved the lowest WERs (10.65% on PFSTAR...)
Wav2Vec2 â evaluatedon â CMU Kids
confidence 100% ¡ Wav2Vec2 achieved the lowest WERs (...22.37% on CMU Kids)
HuBERT â evaluatedon â PFSTAR
confidence 100% ¡ HuBERT yielding similar results on PFSTAR
WavLM â evaluatedon â CMU Kids
confidence 100% ¡ WavLM produced significantly higher error rates (...34.25% on CMU Kids)
Age-specific fine-tuning â improves â Generalization
confidence 90% ¡ fine-tuning on younger childrenâs speech improves generalization to older childrenâs speech
Gender-specific fine-tuning â reduces â Male-preference bias
confidence 90% ¡ fine-tuning reduces male-preference bias
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired articulatory precision. Past research has demonstrated improved recognition through the use of hybrid DNN/HMM sequence discriminative training. This paper presents a comprehensive investigation of various combinations of acoustic features tailored to different Acoustic Models, offering suitable feature selections for each. The incorporation of Pitch features notably improved recognition performance, especially for sentence recognition tasks involving dysarthric speech. Through a systematic examination of the TORGO database, we have demonstrated the potential to enhance the performance of the state-of-the-art Factorized Time Delay Neural Network (F-TDNN) model for recognizing dysarthric speech. Our methods, implemented with the F-TDNN model, resulted in a 4.65\% relative improvement in isolated word recognition and a 4.63\% relative improvement in sentence recognition for dysarthric speech, compared to previous research. This improvement effectively compensates for speech variability, attributable to our deliberate selection of the number of overlapping frames between consecutive training example chunks.
Tags
Links
- Source: https://arxiv.org/abs/2606.19791v1
- Canonical: https://arxiv.org/abs/2606.19791v1
Trouble viewing inline? Open PDF directly â
Full Text
28,045 characters extracted from source content.
Expand or collapse full text
Cross-Dataset, Age, and Gender Generalization: A Comprehensive Analysis of Fine-Tuning Strategies for Low-Resource Childrenâs ASR Abhijit Sinha 1 , Hemant Kumar Kathania 1 , Sudarsana Reddy Kadiri 2 , and Shrikanth Narayanan 2 1 Department of Electronics and Communication Engineering,National Institute of Technology Sikkim, India 2 Signal Analysis and Interpretation Laboratory, University of Southern California, Los Angeles, USA AbstractâChildrenâs speech recognition remains challenging due to acoustic variability, dataset mismatches, and pretraining biases. Self-supervised models like Wav2Vec2 and HuBERT have achieved strong performance in adult ASR, but adapting these models to low-resource childrenâs speech remains limited.In this study, we evaluate age-specific, gender-specific, and cross- dataset fine-tuning strategies on the PFSTAR and CMU Kids datasets. Our findings reveal three key patterns: (i)fine-tuning on younger childrenâs speech improves generalization to older childrenâs speech, (i)fine-tuning reduces male-preference bias, and (i)cross-dataset performance drops significantly due to accent and vocabulary mismatches. Notably, the shorter utterances in the CMU Kids corpus lead to higher baseline WER, highlighting that SSL models struggle with brief speech without adaptation. These insights provide actionable guidelines for developing robust and inclusive ASR systems for children, emphasizing that child- centric ASR benefits from targeted fine-tuning and diverse pretraining data. Index Termsâself-supervised , childrenâs speech recognition, fine-tuning, cross-dataset, low-resource. I. INTRODUCTION Automatic Speech Recognition (ASR) systems have revo- lutionized human-computer interaction, enabling applications such as voice assistants, educational tools, and accessibility services. Despite these advances, accurate ASR for children remains a formidable challenge due to the distinct acoustic and linguistic properties of young speakers. Compared to adult speech, childrenâs speech exhibits a higher pitch, greater variability in pronunciation, and faster speaking rates [1], [2]. These factors contribute to significantly elevated error rates when ASR systems typically trained on extensive adult corpora are applied to children [3]â[5]. Compounding the problem, there is a scarcity of large, labeled datasets for childrenâs speech [6]â[8], which limits the modelsâ ability to capture the full range of variability across different ages and speech patterns. To mitigate these challenges, researchers have explored a range of techniques. Data augmentation methods such as time scale modification [9]â[11], formant modification [12], and vocal tract length normalization [13] have been proposed to artificially increase the diversity of training data. Addition- ally, transfer learning [14], [15] and domain adaptation [16] strategies have been utilized to fine-tune models-originally trained on large adult datasets-with limited child-specific data [17], [18]. More recently, self-supervised learning (SSL)has emerged as a powerful approach for ASR. SSL models such as Wav2Vec2 [19], HuBERT [20], Data2Vec [21], and WavLM [22] can learn robust speech representations from vast amounts of unlabeled audio, and studies have shown that fine-tuning these models on childrenâs speech significantly improves per- formance [16], [23]â[25]. However, most prior work has focused on overall accuracy and has not fully addressed the inherent variability within the child population. ASR generalization across age groups, genders, and datasets remains a key limitation. Younger chil- drenâs speech often exhibits higher acoustic variability,which can impact cross-age performance. Similarly, gender-related differences-such as pitch and articulation-can introducebi- ases from adult pretraining that affect accuracy. Cross-dataset generalization is also a challenge: recognition WER often increases substantially on corpora with different accentsor demographics. Although previous studies showed that fine- tuning on childrenâs speech can reduce some pretraining biases, the specific impacts of age, gender, utterance length, and vocabulary complexity remain under-explored. Motivated by these challenges, our study is the first to systematically analyze age- and gender-specific fine-tuning of SSL models for childrenâs ASR in low-resource scenarios. Specifically, we investigate: â˘How acoustic diversity in younger childrenâs speech influences generalization to older children. â˘Pretraining biases revealed by gender-specific fine-tuning, which may favor male speech. â˘Challenges of cross-dataset generalization, where perfor- mance degrades significantly due to accent and demo- graphic mismatches. We conduct experiments on two well-established childrenâs speech corpora, PFSTAR [17] and CMU Kids [18], using three pre-trained SSL models (Wav2Vec2, HuBERT, and WavLM) under age-specific, gender-specific, and cross-dataset fine-tuning. Our results indicate that models fine-tuned on younger children generalize better to older childrenâs speech, fine-tuning reduces male speech bias, and shorter utterances suffer higher WER, highlighting limitations of current SSL architectures on brief speech. These findings provide practical guidelines for developing robust, child-centric ASR systems. Overall, our study offers insights to guide the design of ASR arXiv:2606.19791v1 [eess.AS] 18 Jun 2026 systems that are both robust and inclusive of childrenâs diverse speech characteristics. I. EXPERIMENTALFRAMEWORK Figure 1 provides a schematic overview of the fine-tuning process for SSL models in the context of childrenâs ASR. The process begins with a pre-trained SSL model such as Wav2Vec2, HuBERT, or WavLM-trained on large, general speech datasets. This model is then fine-tuned on childrenâs speech data, tailored to specific age-group and gender subsets. Fine-tuning adjusts the modelâs parameters so that it better captures the unique characteristics of childrenâs speech,which differ from adult speech in terms of pitch, speaking rate, and pronunciation variability. Pre-Trained SSL Model Latent Feature Encoder Pre-Trained CNN Context Network Pre-Trained Transformer Fine-Tuned SSL Model CTC Tokenizer/ Beam Search Decoder Predicted Text Fine Tuning SSL Model (PF-STAR/CMU Kids) Testing Dataset (PF-STAR/CMU Kids) Training Dataset Overall Age Group & Gender Specific Overall Age Group & Gender Specific Fig. 1. A schematic block diagram illustrating the fine-tuning of SSL models on childrenâs speech data, including both the overall training set and age- group specific subsets, followed by testing on the corresponding test set and its age-group specific subsets. The architecture of these SSL models comprises two main stages. First, a convolutional neural network (CNN) extracts features from raw speech signals, converting them into a sequence of feature vectors. Second, a Transformer-based context network processes these vectors to capture long- range dependencies and temporal relationships in the speech signal. The attention mechanisms within the Transformer allow the model to focus on the most relevant parts of the input sequence, thereby learning contextual patterns and nuances in childrenâs speech. A critical component of SSL is the masking mechanism employed during training. Portions of the input speech features are randomly masked, and the model is tasked with predicting the missing information based on the surrounding context. This self-supervised learning strategy enables the construction of robust speech representations without relying on extensive labeled data, which is particularly beneficial in low-resource scenarios involving childrenâs speech. I. DATASETS ANDEXPERIMENTALSETUP A. Datasets This study utilizes two well-known childrenâs speech datasets: PFSTAR [17] and CMU Kids [18]. The PFSTAR dataset consists of British English recordings of childrenaged 4 to 14 years, with 8.3 hours of training speech from 122 speakers and 1.1 hours of testing speech from 60 speakers. In contrast, the CMU Kids dataset contains American English recordings of children reading sentences, with ages ranging from 6 to 11 years. It features 5180 utterances from 76 speakers, of which 70% (6.3 hours) is used for training and the remaining 30% (2.83 hours) for testing. PFSTAR comprises read speech recorded in quiet environ- ments, whereas CMU Kids includes both read and spontaneous utterances with moderate background noise. Notably, PFSTAR utterances are longer (avg. 41.32 sec vs. 6.28 sec in CMU Kids), which explains its higher total duration despite fewer utterances. Figures 2 and 3 illustrates the dataset distributions used for age-wise, gender-wise, and cross-dataset fine-tuning. PFSTAR is divided into two age groups (4-8 and 9-14 years) and categorized by gender, while CMU Kids is split into two age groups (6-8 and 9-11 years) with balanced gender representation. The broader age range in PFSTAR introduces greater variability in speech patterns, whereas the balanced gender distribution in CMU Kids helps mitigate bias in gender- specific fine-tuning. Gender Distribution 442 414 73 56 Male (Train) Female (Train)Male (Test)Female (Test) 0 100 200 300 400 500 TrainingTesting Age Distribution 0 22 0 73 12 83 8 204 39 81 6 289 42 55 16 25 2 17 2 27 0 4567891011121314 Age (Years) 0 100 200 300 Number of Utterances TrainingTesting Fig. 2. Age and gender distribution of the PFSTAR dataset, showing training and testing splits by speaker gender and age. Gender Distribution 1217 2349 529 1085 Male (Train) Female (Train)Male (Test)Female (Test) 0 1000 2000 TrainingTesting Age Distribution 236 110 772 383 1510 665 982 429 00 66 27 67891011 Age (Years) 0 500 1000 1500 Number of Utterances TrainingTesting Fig. 3. Age and gender distribution of the CMU Kids dataset, showing training and testing splits by speaker gender and age. B. Experimental Setup This section describes the SSL models used and the fine- tuning. 1) Self Supervised Learning (SSL) Models:We employed three state-of-the-art SSL models: Wav2Vec2-Large-960h- lv60-self, HuBERT-Large-LS960-ft, and WavLM-Large, here- after referred to as Wav2Vec2, HuBERT, and WavLM, re- spectively. All three models are pre-trained on large-scale unlabeled audio data to learn robust speech representations, making them highly effective for ASR tasks. They share similar architectures with 25 hidden layers and a feature size of 1024: the initial CNN layer extracts features from raw audio, and the subsequent 24 Transformer layers capture contextual information. Specifically, Wav2Vec2 was pre-trained on 60,000 hours of unlabeled data and fine-tuned on 960 hours of labeled data, while HuBERT used the same unlabeled corpus with a masked prediction strategy. WavLM was pre-trained on 94,000 hours from diverse sources and fine-tuned on 960 hours of labeled data. The architecture processes raw audio through a CNN-based feature encoder, then passes the resulting feature vectors to a Transformer-based context network. A key component is the masking mechanism within the Transformer, where segments of the speech features are randomly masked and predicted from surrounding context, facilitating effective self- supervised learning. Each model employs a different loss function: Wav2Vec2 uses contrastive loss, HuBERT utilizes masked prediction loss with cluster assignments, and WavLM adopts noise-robust loss to handle speech variations. Despite these differences, the overall architecture remains consistent, enabling effective and generalizable speech representations. 2) Fine-Tuning SSL Models for ASR:To fine-tune the SSL models on childrenâs speech data, we followed the framework depicted in Figure 1. Fine-tuning was performed on both the overall dataset and on specific subsets defined by age (i.e., 4-8 and 9-14 years for PFSTAR; 6-8 and 9-11 years for CMU Kids) and gender. Each fine-tuned model was evaluated on multiple test sets, including the complete test set and its corresponding age- and gender-specific subsets. For consistency across experiments, we used a fixed learning rate of 1e-4 and a weight decay of 0.005 to prevent overfit- ting. The models were trained with gradient checkpointing to optimize memory usage and improve generalization. A comprehensive vocabulary covering all transcription characters was built, and Connectionist Temporal Classification (CTC) loss was applied to align predicted sequences with the input speech. Evaluations employed greedy search decoding without an external language model to assess the intrinsic performance of the models. IV. RESULTS ANDDISCUSSION This section presents a comprehensive analysis of the ex- perimental results across various settings. Sections IV-A-IV-E cover the baseline (zero-shot) performance, age group-specific and gender-specific fine-tuning outcomes, full-dataset fine- tuning improvements, and cross-dataset evaluations, examin- ing performance trends, underlying factors, and implications for robust childrenâs ASR. A. Baseline Zero-Shot Results The baseline performance of three SSL models, Wav2Vec2, HuBERT, and WavLM was evaluated in a zero-shot setting on the PFSTAR and CMU Kids datasets. As shown in Table I, Wav2Vec2 achieved the lowest WERs (10.65% on PFSTAR and 22.37% on CMU Kids), with HuBERT yielding similar results on PFSTAR (10.67%) but a higher WER (24.24%) on CMU Kids. In contrast, WavLM produced significantly higher error rates (25.42% on PFSTAR and 34.25%), suggest- ing that its pretraining objectives may not optimally capture the acoustic characteristics of childrenâs speech. WER is consistently higher for CMU Kids, likely due to its shorter utterances, suggesting that these SSL models struggle with shorter utterance lengths. TABLE I WER(%)FROM ZERO-SHOT DECODING USING THREESSLMODELS ON THEPFSTARANDCMU KIDS DATASETS. ModelLibrispeech (Clean)PFSTARCMU Kids Wav2Vec21.9010.6522.37 HuBERT1.9010.6724.24 WavLM-25.4234.25 An age-wise breakdown in Table I reveals that younger age groups consistently incur higher WERs than older groups. For PFSTAR, the 4-8 years group shows WERs approximately 5% (Wav2Vec2) to 6.7% (WavLM) higher than the 9-14 years group. Similar trends are observed in CMU Kids, where the 6-8 years group under performs the 9-11 years group by 6.81% (HuBERT) and 10.8% (WavLM). Gender-wise, PFS- TAR shows male subsets outperforming female ones (e.g., an average difference of 3.39% for Wav2Vec2), while CMU Kids exhibits minimal differences. Table I further compares these results to Librispeech (adult speech), where all models achieve near state-of-the-art performance (e.g., 2.1% for Wav2Vec2), confirming that the degradation on childrenâs speech arises from acoustic mismatches. B. Age Group-wise Fine-Tuning Results Table I shows that fine-tuning on younger childrenâs data yields models that generalize better to older age groups. InPF- STAR, models fine-tuned on the 4-8 years group achieve lower WERs on the 9-14 years test set (e.g., HuBERT: 7.13% vs. 8.67%) compared to fine-tuning on the older group. Similarly, in CMU Kids, Wav2Vec2 fine-tuned on the 6-8 years group achieves a WER of 7.47% on the 9-11 years set, whereas fine- tuning on the older group results in a WER of 11.99%. These trends suggest that greater acoustic variability in younger speech enables models to learn more robust representations. Conversely, models fine-tuned on older children generalizeless effectively to younger children. C. Gender-wise Fine-Tuning Results Table IV summarizes the outcomes of gender-specific fine- tuning. In PFSTAR, models fine-tuned on male data generally achieve lower WERs on female test sets (e.g., Wav2Vec2: 8.37% vs. 6.99% average), with an average difference of 2- 3%. In CMU Kids, the gap is smaller (approximately 1-2%), likely due to its balanced gender distribution. Notably, WavLM shows the largest relative improvement when generalized from male to female, while Wav2Vec2 maintains stable performance TABLE I ZERO-SHOTWER(%)BROKEN DOWN BY AGE GROUP AND GENDER-WISE ON THEPFSTARANDCMU KIDS DATASETS. Model PFSTARCMU Kids Age GroupGenderAge GroupGender 4-89-14 Male Female 6-89-11 Male Female Wav2Vec212.437.368.0611.4524.58 17.77 22.40 22.36 HuBERT13.616.918.1611.4027.03 18.57 25.62 23.51 WavLM31.24 19.62 23.77 25.59 37.76 26.96 34.33 34.20 TABLE I WER(%)ACHIEVED BY FINE-TUNINGSSLMODELS ACROSS DIFFERENT AGE GROUPS WITHIN THEPFSTARANDCMU KIDS DATASETS. PFSTAR Dataset Model Training Age Group: 4-8Training Age Group: 9-14 4-8 (Test)9-14 (Test)Average4-8 (Test)9-14 (Test)Average Wav2Vec28.158.198.178.626.627.62 HuBERT7.977.137.558.676.937.80 WavLM8.349.458.897.877.577.72 CMU Kids Dataset Model Training Age Group: 6-8Training Age Group: 9-11 6-8 (Test)9-11 (Test)Average6-8 (Test)9-11 (Test)Average Wav2Vec23.947.475.7011.994.068.02 HuBERT3.358.555.9512.004.088.04 WavLM2.4010.726.5612.334.458.39 TABLE IV WER(%)ACHIEVED BY FINE-TUNINGSSLMODELS ACROSS GENDER GROUPS WITHIN THEPFSTARANDCMU KIDS DATASETS. PFSTAR Dataset Model Training Gender: MaleTraining Gender: Female Male (Test) Female (Test) Average Male (Test) Female (Test)Average Wav2Vec26.0910.658.375.478.506.99 HuBERT6.9410.708.826.8110.188.50 WavLM7.7310.989.367.569.578.57 CMU Kids Dataset Model Training Gender: MaleTraining Gender: Female Male (Test) Female (Test) Average Male (Test) Female (Test)Average Wav2Vec26.827.066.949.344.516.93 HuBERT7.7013.7210.717.892.925.41 WavLM5.549.597.578.283.075.68 across genders. These results suggest that models fine-tuned on male speech perform better across genders, revealing a pretraining bias favoring male speech. However, this bias is reduced when fine-tuning on gender-balanced datasets, such as CMU Kids. D. Fine-Tuning on Entire Dataset Fine-tuning on the entire training dataset yields significant improvements in WER across both datasets, as shown in Table V. Wav2vec2 yeilds the best results on the overall test sets with WERs of 7.70% and 5.43% for the PFSTAR and CMU Kids dataset respectively. However WavLM shows the largest relative improvement when fine-tuned on the overall datasets. When evaluated on the corresponding age and gender- specific test sets, for PFSTAR, the fine-tuned Wav2Vec2 achieves a WER of 6.94% for the 4-8 years age group and 6.24% for the 9-14 years group, while HuBERT and WavLM show comparable gains. In the CMU Kids dataset, Wav2Vec2 achieves 5.85% and 4.52% for the 6-8 and 9-11 years age groups, respectively. Table VI further illustrates that while TABLE V WER(%)ACHIEVED BY FINE-TUNING THREESSLMODELS ON THE PFSTARANDCMU KIDS DATASETS,INCLUDING BASELINE AND RELATIVE IMPROVEMENT(REL. IMP.)OVER THE BASELINE. Model PFSTARCMU Kids Baseline Fine-Tuned Rel. Imp. Baseline Fine-Tuned Rel. Imp. Wav2Vec210.657.7027.722.375.4375.7 HuBERT10.677.8426.524.245.9675.4 WavLM25.428.0868.234.254.9985.4 fine-tuning on the full dataset significantly reduces WER across all groups, age- and gender-specific fine-tuning provides additional benefits for specific subgroups, highlighting the importance of targeted adaptation. TABLE VI WER(%)ACHIEVED BY FINE-TUNINGSSLMODELS ON THEPFSTAR ANDCMU KIDS DATASETS,EVALUATED ON CORRESPONDING AGE AND GENDER TEST SETS. Model PFSTARCMU Kids Age GroupGenderAge GroupGender 4-8 9-14 Male Female 6-8 9-11 Male Female Wav2Vec2 6.946.24 5.168.505.854.525.685.30 HuBERT6.94 6.55 5.578.276.34 5.16 6.005.94 WavLM6.517.23 5.758.315.084.795.384.77 E. Cross Dataset Evaluation Cross-dataset fine-tuning further highlights the challenge of generalizing across datasets with distinct accents, demograph- ics, and recording conditions. Tables VII and VIII summa- rize the performance when models fine-tuned on one dataset are evaluated on the other. When fine-tuned on PFSTAR (British English) and tested on CMU Kids (American English), WavLM and HuBERT show WER increases of approximately 8-10% due to accent, vocabulary mismatches and greater acoustic variability, whereas Wav2Vec2 exhibits relatively smaller increases in WER. Similarly, models fine-tuned on CMU Kids experience notable degradation when tested on PFSTAR. These results shows that models fine-tuned on one dataset struggle to generalize to another, with WER increasing signifi- cantly compared to the baseline due to accent and demographic mismatches. TABLE VII WER(%)ACHIEVED BY FINE-TUNINGSSLMODELS ONPFSTARAND TESTING ONCMU KIDS,FOR OVERALL AND AGE/GENDER-SPECIFIC SUBSETS. Model Testing Sets CMU Kids Overall Age 6-8 Age 9-11 Male Female Wav2Vec2 34.3737.2128.3433.94 34.59 HuBERT37.3739.8332.1536.84 37.65 WavLM46.6149.3042.3346.13 46.81 Overall, our key findings are: â˘Fine-tuning closes the gap:Although zero-shot ASR on childrenâs speech suffers from acoustic mismatches, targeted fine-tuning yields substantial WER reductions. TABLE VIII WER(%)ACHIEVED BY FINE-TUNINGSSLMODELS ONCMU KIDS AND TESTING ONPFSTAR,FOR OVERALL AND AGE/GENDER-SPECIFIC SUBSETS. Model Testing Sets PFSTAR Overall Age 4-8 Age 9-14 Male Female Wav2Vec2 63.7164.9562.8168.02 70.29 HuBERT79.5180.8378.5479.17 79.96 WavLM75.3279.6672.1674.45 76.51 â˘Subgroup adaptation matters:Age-specific fine-tuning delivers the greatest gains for younger speakers, and gender-specific training helps counteract pretraining bi- ases. â˘Cross-corpus brittleness:Transferring between datasets incurs large performance drops due to accent, vocabulary, and demographic mismatches highlighting the need for more diverse, representative pretraining data. V. CONCLUSION We evaluated SSL models for childrenâs ASR using three fine-tuning strategies (age-specific, gender-specific, andcross- dataset) on the PFSTAR and CMU Kids corpora. Our findings highlight challenges in cross-domain generalization, pretrain- ing biases, and dataset mismatches. Fine-tuning on younger childrenâs speech improved generalization to older children- likely because younger speech has greater acoustic variability- whereas models trained on older children generalized poorly to younger speakers. Gender-specific fine-tuning revealed a bias favoring male speech: models trained on male voices performed relatively better, but balanced gender trainingsets mitigated this bias, underscoring the importance of diverse gender representation. Cross-dataset evaluation showed sig- nificant WER degradation due to accent, vocabulary, and recording differences, emphasizing the need for more diverse pretraining data. We also observed that baseline WER was lower for the longer PFSTAR utterances than for the shorter CMU Kids utterances, suggesting that these pre trained ASR models struggle more with short, acoustically variable speech. Overall, these results indicate that child-centric ASR systems benefit from targeted fine-tuning on specific subgroups and from pretraining on diverse, child-relevant data. REFERENCES [1] S. Lee, A. Potamianos, and S. Narayanan, âAcoustics of childrenâs speech: Developmental changes of temporal and spectral parameters,â The Journal of the Acoustical Society of America, vol. 105, no. 3, p. 1455â1468, 1999. [2] H. K. Vorperian and R. D. Kent, âVowel acoustic space development in children: a synthesis of acoustic and anatomic data.âJournal of speech, language, and hearing research : JSLHR, vol. 50 6, p. 1510â45, 2007. [3] L. L. Koenig, J. C. Lucero, and E. Perlman, âSpeech production variability in fricatives of children and adults: Results of functional data analysis,âThe Journal of the Acoustical Society of America, vol. 124, no. 5, p. 3158â3170, 2008. [4] T. Tran, M. Tinkler, G. Yeung, A. Alwan, and M. Ostendorf,âAnalysis of disfluency in childrenâs speech,âInterspeech, 2020. [5] G. Yeung and A. Alwan, âOn the difficulties of automatic speech recognition for kindergarten-aged children,âInterspeech, 2018. [6] F. Claus, H. G. Rosales, R. Petrick, H.-U. Hain, and R. Hoffmann, âA survey about databases of childrenâs speech.â inInterspeech, 2013, p. 2410â2414. [7] S. Feng, B. M. Halpern, O. Kudina, and O. Scharenborg, âTowards inclusive automatic speech recognition,âComputer Speech & Language, vol. 84, p. 101567, 2024. [8] V. N. Sukhadia and S. A. Chowdhury, âChildrenâs speech recognition through discrete token enhancement,âInterspeech, 2024. [9] Z. Fan, X. Cao, G. Salvi, and T. Svendsen, âUsing modified adult speech as data augmentation for child speech recognition,âIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1â5, 2023. [10] S. Shahnawazuddin, V. Kumar, A. Kumar, and W. Ahmad, âImproving the performance of zero-resource childrenâs asr system through formant and duration modification based data augmentation,âIEEE International Conference on Signal Processing and Communications (SPCOM), p. 1â5, 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:251322654 [11] A. Sinha, M. Singh, S. R. Kadiri, M. Kurimo, and H. K. Kathania, âEffect of speech modification on wav2vec2 models for children speech recognition,â inInternational Conference on Signal Processing and Communications (SPCOM). IEEE, 2024, p. 1â5. [12] H. Kumar Kathania, S. Reddy Kadiri, P. Alku, and M. Kurimo, âStudy of formant modification for children asr,â inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, p. 7429â 7433. [13] T. B. Patel and O. Scharenborg, âImproving end-to-end models for childrenâs speech recognition,âApplied Sciences, 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:268367194 [14] T. Rolland, A. Abad, C. Cucchiarini, and H. Strik, âMultilingual transfer learning for children automatic speech recognition,â inInternational Conference on Language Resources and Evaluation, 2022. [15] J. Thienpondt and K. Demuynck, âTransfer learning for robust low- resource childrenâs speech asr with transformers and source-filter warp- ing,â inInterspeech, 2022. [16] R. Fan and A. Alwan, âDraft: A novel framework to reduce domain shifting in self-supervised learning and its application to childrenâs asr,â inInterspeech, 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:249712229 [17] M. Russell, âThe pf-star british english childrens speech corpus,âThe Speech Ark Limited, 2006. [18] M. Eskenazi, J. Mostow, and D. Graff, âThe cmu kids corpus,âLinguistic Data Consortium, vol. 11, 1997. [19] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, âwav2vec 2.0: A framework for self-supervised learning of speech representations,â Advances in neural information processing systems, vol. 33, p. 12 449â 12 460, 2020. [20] W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, âHubert: Self-supervised speech representation learning by masked prediction of hidden units,âIEEE/ACM transactions on audio, speech, and language processing, vol. 29, p. 3451â3460, 2021. [21] A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, âData2vec: A general framework for self-supervised learning in speech, vision and language,â inInternational Conference on Machine Learning. PMLR, 2022, p. 1298â1312. [22] S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiaoet al., âWavlm: Large-scale self-supervised pre- training for full stack speech processing,âIEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, p. 1505â1518, 2022. [23] R. Fan, Y. Zhu, J. Wang, and A. Alwan, âTowards better domain adaptation for self-supervised models: A case study of child asr,âIEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, p. 1242â1252, 2022. [24] R. Jain, A. Barcovschi, M. Y. Yiwere, D. Bigioi, P. Corcoran, and H. Cucu, âA wav2vec2-based experimental study on self-supervised learning methods to improve child speech recognition,âIEEE Access, vol. 11, p. 46 938â46 948, 2023. [25] J. Li, M. A. Hasegawa-Johnson, and N. L. McElwain, âAnalysis of self- supervised speech models on childrenâs speech and infant vocalizations,â IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW), p. 550â554, 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:267627009