Paper deep dive
Training-Free Intelligibility-Guided Observation Addition for Noisy ASR
Haoyang Li, Changsong Liu, Wei Rao, Hao Shi, Sakriani Sakti, Eng Siong Chng
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/20/2026, 1:36:59 PM
Summary
This paper proposes a training-free, intelligibility-guided Observation Addition (OA) method to improve Automatic Speech Recognition (ASR) performance in noisy environments. By fusing noisy speech and speech-enhanced (SE) signals using weights derived from backend ASR confidence scores, the method avoids the need for training neural predictors or modifying model parameters. Experiments across diverse SE-ASR combinations and datasets (VoiceBank-DEMAND, CHiME-4) demonstrate that this approach outperforms existing OA baselines and provides robust noise-robustness.
Entities (13)
Relation Signals (8)
Conf-OA → istrainingfree → true
confidence 98% · the proposed method is training-free
Demucs → istype → Speech Enhancement Model
confidence 95% · For time-domain SE, we use the causal Demucs
Whisper → istype → ASR System
confidence 95% · Whisper-large ... are employed as strong, noise-robust ASR models
Conf-OA → uses → ASR Confidence
confidence 95% · fusion weights are derived from intelligibility estimates obtained directly from the backend ASR
Conf-OA → outperforms → Classifier-OA
confidence 90% · Conf-OA ... yield more robust performance
Conf-OA → outperforms → DNSMOS-OA
confidence 90% · Conf-OA still outperforms other baselines
CHiME-4 → usedforevaluation → Conf-OA
confidence 90% · we apply SE models ... to Channel 5 of the CHiME-4 test set
VoiceBank-DEMAND → usedfortraining → SE Models
confidence 90% · SE models are trained on the training split of VoiceBank-DEMAND
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Automatic speech recognition (ASR) degrades severely in noisy environments. Although speech enhancement (SE) front-ends effectively suppress background noise, they often introduce artifacts that harm recognition. Observation addition (OA) addressed this issue by fusing noisy and SE enhanced speech, improving recognition without modifying the parameters of the SE or ASR models. This paper proposes an intelligibility-guided OA method, where fusion weights are derived from intelligibility estimates obtained directly from the backend ASR. Unlike prior OA methods based on trained neural predictors, the proposed method is training-free, reducing complexity and enhances generalization. Extensive experiments across diverse SE-ASR combinations and datasets demonstrate strong robustness and improvements over existing OA baselines. Additional analyses of intelligibility-guided switching-based alternatives and frame versus utterance-level OA further validate the proposed design.
Tags
Links
- Source: https://arxiv.org/abs/2602.20967v2
- Canonical: https://arxiv.org/abs/2602.20967v2
Trouble viewing inline? Open PDF directly →
Full Text
28,891 characters extracted from source content.
Expand or collapse full text
Training-Free Intelligibility-Guided Observation Addition for Noisy ASR Haoyang Li 1 , Changsong Liu 1 , Wei Rao 1 , Hao Shi 2 , Sakriani Sakti 3,∗ , Eng Siong Chng 1,∗ 1 Nanyang Technological University, Singapore 2 Independent Researcher 3 Nara Institute of Science and Technology, Japan li0078ng@e.ntu.edu.sg Abstract Automatic speech recognition (ASR) degrades severely in noisy environments. Although speech enhancement (SE) front-ends effectively suppress background noise, they often introduce ar- tifacts that harm recognition. Observation addition (OA) ad- dressed this issue by fusing noisy and SE enhanced speech, im- proving recognition without modifying the parameters of the SE or ASR models. This paper proposes an intelligibility-guided OA method, where fusion weights are derived from intelligi- bility estimates obtained directly from the backend ASR. Un- like prior OA methods based on trained neural predictors, the proposed method is training-free, reducing complexity and en- hances generalization. Extensive experiments across diverse SE-ASR combinations and datasets demonstrate strong robust- ness and improvements over existing OA baselines. Additional analyses of intelligibility-guided switching-based alternatives and frame versus utterance-level OA further validate the pro- posed design. Index Terms:noise-robust automatic speech recognition, speech enhancement post-processing, observation addition 1. Introduction Automatic speech recognition (ASR) aims to transform spo- ken audio signals into their corresponding textual transcriptions. The performance of ASR degrades significantly in the presence of background noise [1–3]. Speech enhancement (SE), which suppresses background noise to estimate cleaner speech [4–7], is widely adopted as a front-end preprocessing step for noise- robust ASR [8–10]. Despite effectively suppressing background noise, SE can introduce artifact errors that negatively impact downstream ASR [11, 12]. This trade-off means SE may de- grade recognition performance, particularly for ASR models al- ready trained for noise robustness. Prior works [13–16] have explored joint training of SE and ASR systems to align the enhancement objective with recogni- tion performance, as SE is typically optimized for signal-level metrics that do not correlate strongly with intelligibility [17]. However, joint training increases computational cost and is in- feasible when the SE and ASR cannot be integrated or are un- known. Moreover, it introduces a trade-off: optimizing for ASR may degrade perceptual speech quality [18] of the SE system. Early studies [11,18] have identified SE-induced artifacts as a primary cause of ASR performance degradation, and demon- strated that Observation Addition (OA) can effectively mitigate such artifacts and improve recognition accuracy. OA is a sim- ple SE post-processing method that interpolates a noisy speech signal y and its enhanced version ˆx = SE(y) along the time * These authors contributed equally as senior authors. dimension using a weighting coefficient S ′ ∈ [0, 1], where ̄x is the fused speech by OA and× denotes multiplication (Eq. 1). In contrast to joint SE-ASR training approaches, OA requires no modification to the underlying SE or ASR models and en- ables explicit control over the trade-off between enhancement strength and recognition performance by adjusting S ′ . ̄x = S ′ × y + (1− S ′ )× ˆx(1) The effectiveness of OA depends critically on the weight- ing coefficient S ′ , which should be adaptively determined to accommodate varying acoustic conditions.Previous works [19–21] use predicted normalized signal quality scores (e.g. DNSMOS [22]) of the noisy input y as S ′ . This reduces the weight of the enhanced signal ˆx when y is relatively clean, and vice versa. However, these approaches do not account for the severity of SE-induced artifacts in ˆx, which may bias S ′ toward the enhanced signal even when substantial artifacts are present, or conversely underutilize ˆx when enhancement is reliable. Re- cent works [21, 23, 24] use neural predictors to estimate S ′ , trained with labeled data derived from recognition metrics such as CER or WER. However, these approaches require ground- truth transcriptions to compute CER/WER, which are often un- available in real-world noisy scenarios. Furthermore, labeling data via ASR and training an additional neural-based predictor increases engineering complexity and may introduce general- ization issues across datasets, SE and ASR systems. In this work, we present a simple yet effective OA frame- work that balances S ′ using speech intelligibility-based met- ric scores computed from y and ˆx, which are directly obtained from backend ASR systems at inference time. By avoiding the training of a dedicated neural-based S ′ predictor, the pro- posed method significantly reduces system complexity and im- proves practicality over prior approaches. Despite its simplicity, the proposed OA method outperforms existing OA approaches across extensive experiments spanning diverse SE models, ASR systems, and datasets. Comparative evaluations against alterna- tive Switch-based and frame-level strategies further validate the effectiveness of the proposed design, establishing the proposed OA as a convenient and broadly applicable SE post-processing method for improving ASR in noisy conditions. 2. Methodology 2.1. Intelligibility-Guided Observation Addition We aim to design a robust OA coefficient S ′ to combine the noisy speech y and enhanced speech ˆx, guided by two de- sign considerations. (1) S ′ should depend on both y and ˆx to adaptively balance their complementary information, and (2) S ′ should be driven by speech intelligibility rather than signal-level arXiv:2602.20967v2 [eess.AS] 6 Jun 2026 Figure 1: The proposed Confidence-guided OA pipeline. quality, as the goal is to improve ASR performance. CER and WER measured by an ASR system serve as direct indicators of speech intelligibility. In an idealized setting where the WERs of both the noisy speech y and the enhanced speech ˆx are available, an intelligibility-based weighting factor can be constructed by normalizing their inverse error rates: S ′ = 1/WER(y) 1/WER(y) + 1/WER(ˆx) .(2) This formulation assigns higher weight to the signal with lower expected recognition error while constraining S ′ ∈ [0, 1]. For numerical stability, a small constant ε = 1× 10 −8 is added to the WERs to prevent division by zero. In practical scenarios, ground-truth transcriptions are un- available and WER cannot be directly computed. We therefore adopt ASR confidence as a practical approximation of speech intelligibility, as it is estimated directly by the ASR system and reflects the model’s internal uncertainty in recognition. Figure 1 illustrates the proposed confidence-based OA framework. Dur- ing inference, confidence scores for both the noisy speech y and the enhanced speech ˆx are directly computed by the frozen backend ASR system. The OA coefficient S ′ is then defined as: S ′ = conf(y) conf(y) + conf(ˆx) ,(3) Likewise, ε is added to prevent zero division. The final output is obtained via Eq. (1) and decoded by the backend ASR. The conf() computation varies by ASR. We consider three popular ASR systems in this study: Whisper [25], Parakeet [26, 27] and Wav2Vec2-CTC [28]. These ASRs are diverse, offering representative examples while keeping the framework general. Whisper outputs, by default, an average log-probability ̄ ℓ k for each decoded segment. We leverage this decoding statis- tic to derive an utterance-level confidence. Specifically, we ex- ponentiate ̄ ℓ k to obtain the segment’s geometric mean token probability. To account for segments of different lengths, the utterance-level confidence is computed as a token-weighted av- erage over all K segments: conf(x) = P K k=1 T k exp( ̄ ℓ k ) P K k=1 T k ,(4) where T k is the number of tokens in segment k. For Parakeet and Wav2Vec2-CTC, token confidence C n is derived from the posterior distribution using Tsallis entropy (q = 0.33), followed by exponential normalization.The utterance-level confidence is defined as the geometric mean over all N token confidences: conf(x) = exp 1 N N X n=1 logC n ! .(5) In Parakeet, token confidence C n is computed directly at each decoding step. whereas in Wav2Vec2-CTC, frame-level confidences are aggregated into token confidence via min pool- ing over greedy CTC spans. 2.2. Intelligibility-Guided Switching As a simpler alternative to OA, we consider a hard switching strategy guided by intelligibility score, in which only the signal with higher ASR confidence is chosen: ̄x = ( y, conf(y)≥ conf(ˆx), ˆx, otherwise. (6) This discrete switching approach serves as a straightforward al- ternative to compare with OA’s weighted combination method. 2.3. Frame-level Observation Addition While prior works [20,21,23,24], Sec. 2.1 and 2.2 mainly focus on utterance-level OA, where S ′ is a single scalar applied uni- formly across all frames, we also consider a more fine-grained frame-level OA formulation in which S ′ is defined as a vector of frame-wise OA coefficients. Specifically, frame-level confi- dence is obtained using the pretrained Wav2Vec2-CTC model. Due to the fixed convolutional strides of Wav2Vec2, each en- coder frame corresponds to a constant number of input sam- ples. This ensures that frame indices of y and ˆx are temporally aligned and have equal duration, which makes frame-wise in- terpolation well defined. Similar to Sec. 2.1, we compute per- frame confidence S ′ t by applying an exponentially normalized transformation to the Tsallis entropy of the posterior distribu- tion, yielding an OA coefficient via Eq. (3) at the frame level. 3. Experiments 3.1. Datasets We evaluate the proposed OA method under both in-domain and out-of-domain conditions using two datasets, with all au- dio sampled at 16 kHz. For in-domain experiments, SE models are trained on the training split of VoiceBank-DEMAND [29], and OA is evaluated on the corresponding test set. VoiceBank- DEMAND is a widely adopted SE benchmark that provides paired clean speech and synthetically corrupted noisy signals. Although drawn from the same domain, the test set includes speakers, noise types and SNR levels not seen during training, introducing a controlled mismatch. For out-of-domain evaluation, we apply SE models trained on VoiceBank-DEMAND to Channel 5 of the CHiME-4 test set [30]. CHiME-4 is a widely used benchmark for distant- talking speech recognition, comprising recordings captured by a multi-microphone tablet array in real-world noisy environ- ments. The selected test partition contains 1,320 real noisy ut- terances recorded across four everyday acoustic scenes: bus, cafe, pedestrian area, and street junction, along with 1,320 cor- responding synthetically corrupted noisy signals. This out-of- domain setting reflects realistic deployment scenarios, where Table 1: WER of different utterance-level post-processing methods. ˆx is obtained by GR-KAN-based MPSENet. Method Voicebank+DemandCHIME-4 SimuCHIME-4 Real WhisperParakeetWav2Vec2WhisperParakeetWav2Vec2WhisperParakeetWav2Vec2 Noisy y3.122.1811.395.365.2425.826.486.2242.24 Enhanced ˆx2.341.528.098.007.9120.7015.7514.7230.91 SNR-OA no-clip [20]2.401.798.01– SNR-OA clip [20]2.511.699.12– DNSMOS-OA [21] 2.361.607.885.275.0017.0911.569.7526.26 Classifier-OA 2class [23]2.311.307.717.126.7618.5912.8811.8727.73 Classifier-OA 3class [23] 2.371.578.285.064.8816.906.185.3724.85 Conf-Switch (Eq. 6)2.011.267.645.935.3219.487.556.0728.58 Conf-OA (Eq. 3) 2.241.357.614.974.7816.735.865.5524.03 WER-OA (Eq. 2)1.551.036.564.484.1915.505.364.8523.43 Table 2: WER of different utterance-level post-processing methods. ˆx is obtained by Demucs. Method Voicebank+DemandCHIME-4 SimuCHIME-4 Real WhisperParakeetWav2Vec2WhisperParakeetWav2Vec2WhisperParakeetWav2Vec2 Noisy y3.122.1811.395.365.2425.826.486.2242.24 Enhanced ˆx 3.912.6410.6328.6927.4248.0729.6527.5858.99 SNR-OA no-clip [20]3.682.699.88– SNR-OA clip [20] 2.941.9210.21– DNSMOS-OA [21]3.122.439.936.826.6625.2017.7516.3345.97 Classifier-OA 2class [23] 3.432.5210.018.167.3528.3110.9010.8740.21 Classifier-OA 3class [23]2.762.109.775.705.4823.086.866.6136.01 Conf-Switch (Eq. 6)3.022.229.755.765.5825.767.236.2641.65 Conf-OA (Eq. 3)2.762.119.695.645.4022.876.606.1035.84 WER-OA (Eq. 2)2.361.718.625.054.9421.915.975.7434.96 Table 3: WER of OA (Sec.2.1) and Switch (Sec.2.2) on CHIME4, ˆx is obtained by GR-KAN-based MP-SENet. ASR is Parakeet. Method Confidence-CorrectMiscalibratedAmbiguous OA WinSwitch WinTieOA WinSwitch WinTieOA WinSwitch WinTie # Samples38(4.42%)67(7.80%)754(87.78%)130(54.17%)4(1.67%)106(44.17%)42(2.73%)21(1.36%)1478(95.91%) Noisy y21.1912.166.669.5822.0213.4019.127.253.24 Enhanced ˆx24.7714.5324.9812.0626.9716.1119.127.253.24 Conf-Switch (Eq. 6) 16.237.165.7916.0231.3915.2519.127.253.24 Conf-OA (Eq. 3)5.9716.525.795.2540.5015.258.1516.203.24 OA is performed using SE models trained under conditions that differ substantially from those encountered at test time. 3.2. Models 3.2.1. SE We employ both time-domain and time-frequency (TF) domain SE systems, representing the two main categories in SE. For time-domain SE, we use the causal Demucs [4], a 1D CNN and LSTM-based model. For TF-domain SE, we use the GR- KAN–based MP-SENet from [31], a state-of-the-art non-causal SE. The two SE systems cover different architectures, input do- mains, causalities, and performance levels. 3.2.2. ASR We evaluate the proposed OA method using three ASR systems with diverse architectures and robustness levels. Whisper-large [25] and Parakeet [26, 27] are employed as strong, noise-robust ASR models, representing large-scale sequence-to-sequence and TDT-based transducer frameworks, respectively. For Whis- per, we use the Whisper-large model, while for Parakeet we adopt the parakeet-tdt-0.6b-v2 model 1 . In addition, we include a wav2vec2-large ASR model [28] fine-tuned on the full Lib- riSpeech [32] labeled dataset 2 , which serves as a comparatively weaker and more noise-sensitive CTC-based baseline. This setup allows us to evaluate OA under contrasting conditions where either noisy or enhanced speech may be preferred. 3.3. Implementation 3.3.1. SE training For Demucs, we use the causal version with depth 5, hidden dimension of 48 at depth 1. Kernel, stride and resample factor are set to 8, 4, 4, respectively. Training uses batch 16, LR 3e-4, adam optimizer for 500 epochs. For GR-KAN MP-SENet, we use the same model configuration as in [31], and trained for 200 1 https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2 2 https://huggingface.co/facebook/wav2vec2-large-960h epochs at batch 4. Optimizer uses AdamW (β1 = 0.8 and β2 = 0.99). LR is set to 5e-4 and decays by 0.99 every epoch. 3.4. Baselines We compare the proposed methods with OA solutions based on speech quality [20, 21] and intelligibility [23], which share the OA formulation in Eq. 1, differing only in how S ′ is computed. SNR-OA clip and SNR-OA no-clip [20]. S ′ is defined as the normalized SNR of the noisy speech y, based on the observa- tion that ASR performance on y typically improves at higher SNRs. Although the original method uses a neural SNR esti- mator to handle unknown SNRs in practice, we directly use the ground-truth SNR available in VoiceBank-DEMAND to evalu- ate the method’s upper-bound performance without confound- ing estimation errors. Following the original work, S ′ is clipped to [0.6, 1] (SNR-OA clip ). we also report results without clipping (SNR-OA no-clip ) for fair comparison with other OA methods that do not apply this constraint. DNSMOS-OA [21]. S ′ is defined as the average of the normalized BAK and SIG scores of y, where BAK and SIG are DNSMOS components for background noise and signal quality. As in the SNR-based method, we directly use DNSMOS scores without training a separate predictor. Classifier-OA 2class [23] and Classifier-OA 3class . In [23], a neural-based binary classifier is trained with labels defined as class = ( 0, d(x t ,y t ) < d(x t , ˆx t ), 1, otherwise, (7) where d(·,·) denotes the edit distance, x t , y t , ˆx t are the transcript of ground-truth, noisy speech and enhanced speech, respectively. S ′ is defined as ˆp 0 , the posterior probabili- ties of class 0. However, this binary formulation, denoted by Classifier-OA 2class , assigns tie cases to class 1, intro- ducing bias.Therefore, we introduced a 3-class variant (Classifier-OA 3class ), and S ′ is equivalent to ˆp 0 + ˆp 2 × 0.5: class = 0, d(x t ,y t ) < d(x t , ˆx t ), 1, d(x t ,y t ) > d(x t , ˆx t ), 2, d(x t ,y t ) == d(x t , ˆx t ), (8) The model follows [23] and is trained on VoiceBank- DEMAND, with enhanced speech generated by GR-KAN MP- SENet (Table 1) and Demucs (Table 2). Training labels come from Whisper-large, with WER as the edit distance d(·,·). 4. Results and discussion 4.1. Comparison with previous OA methods Tables 1 and 2 compare various OA methods. Bold/underlined values denote best/second-best results, while italicized values require WER/ground-truth text access. WER-OA (Eq. 2) con- sistently achieves the lowest WER in all test cases, support- ing the design of intelligibility-guided OA in Section 2.1. The proposed Conf-OA (Eq. 3) achieves the best overall practical performance, confirming ASR confidence as a reliable alterna- tive in the OA task. We note three cases in Table 2 (CHiME- 4 Simu+Whisper, CHiME-4 Simu+Parakeet, and CHiME-4 Real+Whisper) where Conf-OA slightly exceeds the WER of the noisy speech. This occurs when the relative performance gap between noisy and enhanced speech is large, making the weaker signal insufficient to improve the stronger one. Nev- ertheless, Conf-OA still outperforms other baselines, demon- strating greater robustness. We also observe that our 3-class classifier variant (Classifier-OA 3class ) generally outperforms the original 2-class classifier (Classifier-OA 2class ). However, it still requires training an explicit model, unlike the proposed training-free Conf-OA approach, which is much simpler by de- sign and yield more robust performance. 4.2. Analysis of OA and Switching Table 3 further analyzes the behavior of Conf-OA (Eq. 3) ver- sus Conf-Switch (Eq. 6). Evaluation is performed on CHiME- 4 Simu+Real, using GR-KAN MP-SENet for SE and Para- keet for ASR. Test data are grouped as: 1) Ambiguous: noisy and enhanced speech have equal WER; 2) Confidence-Correct: higher-WER speech has lower confidence; 3) Miscalibrated: higher-WER speech has higher confidence. Each group is fur- ther split by whether OA outperforms Switch or ties. The benefit of OA is most evident in the miscalibrated cases, where OA equals or outperforms Switch in nearly all in- stances. Notably, under the OA Win sub-group which accounts for more than half of all miscalibrated cases, OA generally im- proves WER over both the noisy and enhanced speech, whereas Switch generally degrades recognition. This is because, under miscalibration, Switch always selects the higher-WER speech, while OA can partially compensate for the weaker signal by leveraging the stronger one. Another observation from the am- biguous cases is that OA rarely affects recognition when the noisy and enhanced speech have identical WER. 4.3. Does performance improve with frame-level OA? Table 4: WER of utterance- (Section 2.1) and frame-level OA (Section 2.3) on Chime4 Simu/Real partition. ASR is Wav2Vec2. Method GR-KAN MP-SENetDemucs SimuRealSimuReal Utterance-level OA16.7324.0322.8735.84 Frame-level OA16.8725.3025.1937.76 Table 4 compares utterance-level (Section 2.1) and frame- level (Section 2.3) using Conf-OA (Eq. 3). We observe perfor- mance degradation with frame-level OA. This is likely because frame-level fusion introduces inconsistent weighting across ad- jacent frames, disrupting the temporal continuity expected by the ASR. In contrast, utterance-level OA preserves global con- sistency, resulting in more stable recognition performance. 5. Conclusion This work proposes an intelligibility-guided SE post-processing framework for noise-robust ASR. By performing OA with fu- sion weights derived from backend ASR confidence scores, the proposed method effectively balances noisy and enhanced speech, leading to recognition improvements across multiple SE-ASR systems and datasets, and outperforming existing OA approaches. Importantly, the method operates in an inference- only manner, avoiding the additional training stage required by prior neural-based OA methods, reducing system complexity. Further analyses of switching-based and frame-level variants provide additional evidence for the effectiveness of the pro- posed confidence-based, utterance-level OA strategy. Overall, this work provides an effective, easy-to-implement, and broadly applicable solution for improving ASR robustness in noisy con- ditions without modifying existing SE or ASR models. 6. Generative AI Use Disclosure Generative AI tools were used for limited editorial support, in- cluding grammar checking, redundant text removal, and assis- tance with LaTeX equations. All scientific work, such as the introduction, methodology, experiments, results, and conclu- sions, was conducted by the authors. All authors reviewed the manuscript and take full responsibility for the final submission. 7. References [1] T. Virtanen, R. Singh, and B. Raj, Techniques for noise robustness in automatic speech recognition. John Wiley & Sons, 2012. [2] J. Li, L. Deng, Y. Gong, and R. Haeb-Umbach, “An overview of noise-robust automatic speech recognition,” IEEE/ACM Transac- tions on Audio, Speech, and Language Processing, vol. 22, no. 4, p. 745–777, 2014. [3] A. Rodrigues, R. Santos, J. Abreu, P. Bec ̧a, P. Almeida, and S. Fer- nandes, “Analyzing the performance of asr systems: The effects of noise, distance to the device, age and gender,” in Proceedings of the X International Conference on Human Computer Interac- tion, 2019, p. 1–8. [4] A. Defossez, G. Synnaeve, and Y. Adi, “Real time speech enhancement in the waveform domain,”arXiv preprint arXiv:2006.12847, 2020. [5] Y.-X. Lu, Y. Ai, and Z.-H. Ling, “Mp-senet: A speech enhance- ment model with parallel denoising of magnitude and phase spec- tra,” arXiv preprint arXiv:2305.13686, 2023. [6] H. Li, J. Q. Yip, T. Fan, and E. S. Chng, “Speech enhancement using continuous embeddings of neural audio codec,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, p. 1–5. [7] H. Li, N. Hou, Y. Hu, J. Yao, S. M. Siniscalchi, and E. S. Chng, “Aligning generative speech enhancement with human preferences via direct preference optimization,” arXiv preprint arXiv:2507.09929, 2025. [8] M. Delcroix, T. Yoshioka, A. Ogawa, Y. Kubo, M. Fujimoto, N. Ito, K. Kinoshita, M. Espi, S. Araki, T. Hori et al., “Strate- gies for distant speech recognitionin reverberant environments,” EURASIP Journal on Advances in Signal Processing, vol. 2015, no. 1, p. 60, 2015. [9] A. Nicolson and K. K. Paliwal, “Deep xi as a front-end for robust automatic speech recognition,” in 2020 IEEE Asia-Pacific Confer- ence on Computer Science and Data Engineering (CSDE). IEEE, 2020, p. 1–6. [10] Y. Yang, A. Pandey, and D. Wang, “Towards decoupling frontend enhancement and backend recognition in monaural robust asr,” Computer Speech & Language, vol. 95, p. 101821, 2026. [11] K. Iwamoto, T. Ochiai, M. Delcroix, R. Ikeshita, H. Sato, S. Araki, and S. Katagiri, “How bad are artifacts?: Analyzing the impact of speech enhancement errors on asr,” arXiv preprint arXiv:2201.06685, 2022. [12] Y. Hu, N. Hou, C. Chen, and E. S. Chng, “Interactive feature fu- sion for end-to-end noise-robust speech recognition,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, p. 6292–6296. [13] Z.-Q. Wang and D. Wang, “A joint training framework for ro- bust automatic speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 24, no. 4, p. 796– 806, 2016. [14] T. Menne, R. Schl ̈ uter, and H. Ney, “Investigation into joint op- timization of single channel speech enhancement and acoustic modeling for robust asr,” in ICASSP 2019-2019 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019, p. 6660–6664. [15] Y. Koizumi, S. Karita, A. Narayanan, S. Panchapagesan, and M. Bacchiani, “Snri target training for joint speech enhancement and recognition,” arXiv preprint arXiv:2111.00764, 2021. [16] D. Ma, N. Hou, H. Xu, E. S. Chng et al., “Multitask-based joint learning approach to robust asr for radio communication speech,” in 2021 Asia-Pacific Signal and Information Processing Associa- tion Annual Summit and Conference (APSIPA ASC). IEEE, 2021, p. 497–502. [17] S.-J. Chen, A. S. Subramanian, H. Xu, and S. Watanabe, “Build- ing state-of-the-art distant speech recognition using the chime-4 challenge with a setup of speech enhancement baseline,” arXiv preprint arXiv:1803.10109, 2018. [18] K. Iwamoto, T. Ochiai, M. Delcroix, R. Ikeshita, H. Sato, S. Araki, and S. Katagiri, “How does end-to-end speech recog- nition training impact speech enhancement artifacts?” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).IEEE, 2024, p. 11 031– 11 035. [19] Y.-W. Chen, J. Hirschberg, and Y. Tsao, “Noise robust speech emotion recognition with signal-to-noise ratio adapting speech en- hancement,” arXiv preprint arXiv:2309.01164, 2023. [20] K.-C. Wang, Y.-J. Li, W.-L. Chen, Y.-W. Chen, Y.-C. Wang, P.- C. Yeh, C. Zhang, and Y. Tsao, “Bridging the gap: Integrating pre-trained speech enhancement and recognition models for ro- bust speech recognition,” in 2024 32nd European Signal Process- ing Conference (EUSIPCO). IEEE, 2024, p. 426–430. [21] Z. Cui, C. Cui, T. Wang, M. He, H. Shi, M. Ge, C. Gong, L. Wang, and J. Dang, “Reducing the gap between pretrained speech en- hancement and recognition models using a real speech-trained bridging module,” in ICASSP 2025-2025 IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, p. 1–5. [22] C. K. Reddy, V. Gopal, and R. Cutler, “Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise sup- pressors,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).IEEE, 2021, p. 6493–6497. [23] H. Sato, T. Ochiai, M. Delcroix, K. Kinoshita, N. Kamo, and T. Moriya, “Learning to enhance or not: Neural network- based switching of enhanced and observed signals for overlap- ping speech recognition,” in ICASSP 2022-2022 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, p. 6287–6291. [24] S. Huang, Y. Du, J. Yang, D. Zhang, X. Jia, J. Deng, J. Kang, and R. Zheng, “Overlap-adaptive hybrid speaker diarization and asr-aware observation addition for misp 2025 challenge,” arXiv preprint arXiv:2505.22013, 2025. [25] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International conference on machine learning. PMLR, 2023, p. 28 492–28 518. [26] D. Rekesh, N. R. Koluguri, S. Kriman, S. Majumdar, V. Noroozi, H. Huang, O. Hrinchuk, K. Puvvada, A. Kumar, J. Balam et al., “Fast conformer with linearly scalable attention for efficient speech recognition,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023, p. 1–8. [27] H. Xu, F. Jia, S. Majumdar, H. Huang, S. Watanabe, and B. Gins- burg, “Efficient sequence transduction by jointly predicting tokens and durations,” in International Conference on Machine Learn- ing. PMLR, 2023, p. 38 462–38 484. [28] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,” Advances in neural information processing systems, vol. 33, p. 12 449–12 460, 2020. [29] C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Investigating rnn-based speech enhancement methods for noise- robust text-to-speech.” in SSW, 2016, p. 146–152. [30] E. Vincent, S. Watanabe, J. Barker, and R. Marxer, “The 4th chime speech separation and recognition challenge,” URL: http://spandh. dcs. shef. ac. uk/chime challenge/(last accessed on 1 August, 2018), 2016. [31] H. Li, Y. Hu, C. Chen, S. M. Siniscalchi, S. Liu, and E. S. Chng, “From kan to gr-kan: Advancing speech enhancement with kan- based methodology,” arXiv preprint arXiv:2412.17778, 2024. [32] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Lib- rispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2015, p. 5206–5210.