Paper deep dive
PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction
Wanyi Ning, Wei Zhou, Yingpeng Li, Yinshang Guo, Haitao Qian, Yiming Cheng
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/10/2026, 3:58:50 AM
Summary
PS4 is a proxy-supervised joint training framework for target speaker extraction (TSE) that addresses the lack of clean reference speech in real conversational mixtures. It constructs a large-scale corpus (REAL-PS4) from four public datasets and fine-tunes a BSRNN-based model using four differentiable proxy objectives: ASR cross-entropy, speaker similarity, frame-level VAD, and perceptual audio quality (DNSMOS). PS4 ranks 2nd overall on the REAL-T challenge leaderboard, achieving the best speaker similarity and timing F1 scores.
Entities (11)
Relation Signals (8)
PS4 → evaluatedon → REAL-T
confidence 95% · On the REAL-T challenge leaderboard, PS4 ranks 2nd overall
PS4 → trainedon → REAL-PS4
confidence 95% · construct the first large-scale proxy-supervised training corpus, REAL-PS4
PS4 → uses → BSRNN
confidence 95% · fine-tunes a BSRNN-based TSE model
PS4 → uses → ECAPA-TDNN
confidence 95% · comprises a BSRNN separator and a pretrained ECAPA-TDNN speaker encoder
PS4 → achievesbest → Speaker Similarity
confidence 90% · achieving the best speaker similarity and timing F1 among all submitted systems
REAL-PS4 → derivedfrom → AISHELL-4
confidence 90% · construct a large-scale proxy-supervised training corpus, termed REAL-PS4, from four publicly available real conversational speech datasets: AISHELL-4
Proxy-Supervised Joint Training → optimizes → ASR Cross-Entropy Loss
confidence 90% · four complementary differentiable objectives: ASR cross-entropy
Proxy-Supervised Joint Training → optimizes → DNSMOS Loss
confidence 90% · perceptual audio quality... employ a differentiable implementation of DNSMOS as an auxiliary objective
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision are unavailable. We present PS4, a proxy-supervised training framework for TSE in real conversational mixtures, with two main contributions. First, we construct a large-scale corpus of 71,771 training samples derived from four public datasets, covering both Chinese and English scenarios. Each sample contains an overlapping speech mixture, per-speaker enrollment audio, a ground-truth transcript, and frame-level voice activity labels. Second, we propose a proxy-supervised joint training strategy that fine-tunes a BSRNN-based TSE model using four complementary differentiable objectives: ASR cross-entropy, speaker similarity, frame-level voice activity detection, and perceptual audio quality. Starting from a publicly available pre-trained checkpoint, only the BSRNN separator is updated during fine-tuning. On the REAL-T challenge leaderboard, PS4 ranks 2nd overall, achieving the best speaker similarity and timing F1 among all submitted systems.
Tags
Links
- Source: https://arxiv.org/abs/2607.08111v1
- Canonical: https://arxiv.org/abs/2607.08111v1
Trouble viewing inline? Open PDF directly →
Full Text
22,418 characters extracted from source content.
Expand or collapse full text
PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction Wanyi Ning 1,2 , Wei Zhou 1 , Yingpeng Li 1 , Yinshang Guo 3 , Haitao Qian 1 , Yiming Cheng 1 1 Yijiahe AI, Nanjing, China 2 Tianjin University, Tianjin, China 3 Nanjing University, Nanjing, China ningwanyi@126.com Abstract—Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for su- pervision are unavailable. We present PS4, a proxy-supervised training framework for TSE in real conversational mixtures, with two main contributions. First, we construct a large-scale corpus of 71,771 training samples derived from four public datasets, covering both Chinese and English scenarios. Each sample contains an overlapping speech mixture, per-speaker enrollment audio, a ground-truth transcript, and frame-level voice activity labels. Second, we propose a proxy-supervised joint training strategy that fine-tunes a BSRNN-based TSE model using four complementary differentiable objectives: ASR cross- entropy, speaker similarity, frame-level voice activity detection, and perceptual audio quality. Starting from a publicly available pre-trained checkpoint, only the BSRNN separator is updated during fine-tuning. On the REAL-T challenge 1 leaderboard, PS4 ranks 2nd overall, achieving the best speaker similarity and timing F1 among all submitted systems. Index Terms—target speaker extraction, proxy supervision, joint training, real conversational speech I. INTRODUCTION Target speaker extraction (TSE) aims to isolate the speech of a specific speaker from a multi-talker mixture given a short enrollment utterance as a reference [1]–[5]. It has at- tracted increasing interest as a core component of personalized speech interfaces, meeting transcription, and assistive hearing systems. State-of-the-art TSE models [6]–[10] have achieved impressive results on standard benchmarks such as VoxCeleb [11], WSJ0-2mix [12] and LibriMix [13]. However, these benchmarks are constructed by artificially mixing clean single- speaker recordings. In contrast, real conversational recordings exhibit substantially different characteristics: reverberation, background noise, device-specific distortions, and natural turn- taking patterns that produce irregular overlap durations and speaker ratios [14]–[19]. Because clean reference signals for individual speakers are unavailable in real-world recordings, it is not straightforward to apply conventional signal-level supervision like SI-SNR loss [6], [20] to train TSE models on such data. As a result, existing TSE systems are predominantly trained on simulated mixtures and may degrade when deployed in real conversational scenarios. To bridge the gap between simulated benchmarks and real- world deployment, REAL-T [19] was introduced as the first benchmark for evaluating TSE systems on real conversational 1 https://real-tse.github.io/challenge/ mixtures. It provides carefully curated real multi-talker record- ings with corresponding speaker enrollment utterances and evaluation annotations, enabling systematic assessment of TSE models under realistic acoustic conditions. However, REAL- T is solely an evaluation benchmark and does not provide a matched training corpus or clean target speech for supervision. Consequently, existing TSE models, including the official REAL-T baseline, still rely on simulated mixtures for training [1], [2], [9], [19]. How to effectively leverage large-scale real conversational recordings for TSE training therefore remains an open challenge. In this paper, we present PS4, a proxy-supervised training framework for target speaker extraction from real conversa- tional recordings. Instead of relying on unavailable clean target speech, PS4 leverages multiple proxy supervision signals that can be obtained directly from real conversational data. Our main contributions are summarized as follows: • We construct the first large-scale proxy-supervised train- ing corpus, REAL-PS4 2 , by reformatting four public meeting and conversational speech datasets into a unified training format compatible with the REAL-T benchmark [19]. The resulting corpus contains 71,771 training sam- ples covering both Chinese and English scenarios, pro- viding overlapping speech mixtures, speaker enrollment utterances, transcripts, and voice activity annotations. • We propose a proxy-supervised joint optimization frame- work PS4 3 that enables TSE training with multiple complementary differentiable objectives from linguistic, speaker, temporal, and perceptual perspectives while re- maining fully compatible with gradient optimization. • The experiment results on the REAL-T benchmark [19] demonstrate that PS4 significantly improves the official baseline under the challenge evaluation protocol, provid- ing an effective and practical solution for training TSE models directly on real conversational recordings. I. REAL-PS4 CORPUS We construct a large-scale proxy-supervised training cor- pus, termed REAL-PS4, from four publicly available real conversational speech datasets: AISHELL-4 [17], AliMeeting [15], AMI [14], and CHiME-6 [16]. As shown in Fig. 1, we generate enrollment utterances and target-speaker training 2 https://huggingface.co/datasets/TaurenMountain/REAL-PS4 3 Our code is available at https://github.com/TaurenMountain/PS4. arXiv:2607.08111v1 [cs.SD] 9 Jul 2026 Input Datasets Enrollment Extraction Mixture Generation AISHELL-4 (ZH, Meeting) AliMeeting (ZH, Meeting) AMI (EN, Meeting) CHiME-6 (EN, Dinner) Isolate Single-Speaker Regions Filter & Truncate Keep >5s non-silent. Cap at 10s. Enrollment Utterances Mixture Audio Target Transcripts VAD Labels Segment & Truncate Merge <0.5s gaps. Cap at 30s. Identify Overlap Regions (≥2 speakers) Quality Filter PS4 Training Algorithm OUTPUT: REAL-PS4 1. Target > 20% duration. 2. Text ≥ 2 valid chars. BSRNN ECAPA TSE Model Mixture Enrollment + Linguistic Speaker Temporal Perceptual Whisper Large-v3 ResNet34 DNSMOS Frame Energy Cross-Entropy Loss Hinge Ranking Loss BCE Loss OVRL Loss Total Loss ∑ Extracted Speech Proxy-Supervised Backprop Fig. 1. The workflow of PS4, including Corpus Construction and model training. samples through two parallel processing branches, followed by a final quality filtering stage. Enrollment extraction. For each recording session, speaker diarization annotations are first used to identify regions where exactly one speaker is active. These single-speaker segments are treated as candidate enrollment utterances. To ensure sufficient speaker information, only non-silent segments longer than 5 s are retained, and up to five enrollment clips are selected for each speaker within a session. Mixture generation. Mixture samples are constructed from overlapping speech regions where two or more speakers are simultaneously active. Adjacent overlap regions separated by less than 0.5 s are merged to avoid excessive fragmentation, while merged segments shorter than 5 s or longer than 100 s are discarded. For each retained segment, the corresponding mixture waveform is extracted from the original recording and peak-normalized. The target speaker’s transcript is obtained from the original annotations after removing non-speech mark- ers. In addition, frame-level voice activity detection (VAD) [21] labels are generated by projecting the target speaker’s diarization intervals onto the extracted mixture segment, pro- viding temporal supervision during training. Quality filtering. Finally, several quality-control rules are applied to improve the reliability of the corpus. We retain only samples in which the target speaker occupies more than 20% of the mixture duration and the cleaned transcript contains at least two valid characters. Since Whisper [22] accepts at most 30 s of audio as input, mixture segments are truncated to this duration. Enrollment utterances are further limited to 10 s to reduce GPU memory consumption during training. The remaining samples constitute the final REAL-PS4 corpus. I. METHOD: PS4 We propose PS4, a proxy-supervised training framework with four complementary objectives from real conversational recordings, illustrated in Figure 1. Our method leverages supervision signals readily available in REAL-PS4, including transcripts, speaker enrollment, diarization annotations. We build upon the pretrained BSRNN-ECAPA model released in the REAL-T open-source repository 4 [19]. It comprises a BSRNN separator [23] and a pretrained ECAPA-TDNN speaker encoder [24], initialized from a checkpoint pretrained on VoxCeleb1 [11]. During training, we fine-tune only the separator while keeping the speaker encoder frozen. We jointly optimize the separator using four complementary proxy supervision objectives, which constrain the extracted speech from linguistic, speaker, temporal and perceptual per- spectives, respectively. The overall training objective is defined as L = λ ce L CE + λ sim L SIM + λ dns L DNSMOS + λ vad L VAD ,(1) where λ ce , λ sim , λ dns , and λ vad denote the corresponding loss weights. Linguistic supervision. The extracted speech should pre- serve the linguistic content of the target speaker. We therefore employ a frozen Whisper large-v3 [22] model as a differen- tiable teacher and compute the cross-entropy loss [25] between the predicted token distribution and the ground-truth transcript: L CE = CrossEntropy (logits(f Whisper (ˆs)),y text ),(2) where ˆs denotes the extracted waveform and y text is the reference transcript. This objective encourages the separator to produce speech that remains accurately recognizable while allowing gradients to propagate through the differentiable Whisper frontend. Speaker supervision. To preserve the target speaker iden- tity while suppressing interfering speakers, we introduce a speaker similarity objective based on cosine-margin ranking. Specifically, the similarity between the extracted speech and the enrollment utterance is encouraged to exceed that between the original mixture and the enrollment by a predefined margin: L SIM =E [max (0,m− (cos(ˆs,e)− cos(x,e)))],(3) where e denotes the enrollment utterance, x is the input mix- ture, and m is the margin. Speaker embeddings are extracted using frozen ResNet34 5 speaker encoders [26] Temporal supervision. The diarization annotations used during corpus construction naturally provide frame-level target speaker activity labels. We exploit this information as an auxiliary temporal supervision by minimizing the binary cross- entropy between the predicted frame-level speech activity of the extracted waveform and the target VAD labels: L VAD = BCE (σ(Energy(ˆs)),y VAD ),(4) 4 https://w.modelscope.cn/datasets/wenet/wesep pretrainedmodels 5 https://modelscope.cn/datasets/wenet/wespeakerpretrainedmodels AISHELL-4 AliMeeting AMI CHiME-6 DipCo 0.4 0.6 0.8 1.0 TER AISHELL-4 AliMeeting AMI CHiME-6 DipCo 0.4 0.5 0.6 0.7 Speaker SIM AISHELL-4 AliMeeting AMI CHiME-6 DipCo 2 3 DNSMOS OVRL AISHELL-4 AliMeeting AMI CHiME-6 DipCo 0.800 0.825 0.850 0.875 0.900 Timing F1 BSRNN-EMBBSRNN-TFMAPPS4 (ours) Fig. 2. Per-dataset evaluation on the REAL-T development set across four metrics. TABLE I PER-DATASET RESULTS OF PS4 ON THE REAL-T DEVELOPMENT SET. DatasetN TER↓ SIM↑ DNSMOS OVRL↑F1↑ AISHELL-42400.5670.5753.1560.900 AliMeeting4810.4400.6333.3690.861 AMI5920.3880.7143.5490.902 CHiME-65450.5520.5673.3050.891 DipCo1330.4760.6223.3400.897 Overall19910.4730.6313.3770.888 wherey VAD denotes the target speaker activity labels derived from the diarization annotations. This objective encourages the separator to produce speech that is temporally consistent with the target speaker activity. Perceptual supervision. Besides preserving linguistic and speaker information, the extracted speech should also exhibit good perceptual quality. We therefore employ a differentiable implementation of DNSMOS [27] as an auxiliary objective, L DNSMOS =−DNSMOS OVRL (ˆs),(5) where DNSMOS OVRL denotes the overall perceptual quality score predicted by the DNSMOS model 6 . This objective directly encourages higher perceptual speech quality without requiring clean reference signals. IV. EXPERIMENTS We conduct experiments on the REAL-T [19] development set, which contains 1,991 samples from five conversational speech corpora: AISHELL-4 [17], AMI [14], AliMeeting [15], CHiME-6 [16], and DipCo [18]. We follow the official REAL-T evaluation protocol and report four metrics: token error rate (TER), speaker similarity (SIM), DNSMOS OVRL perceptual quality, and timing F1 [19]. TER is minimized, while the remaining three metrics are maximized. We com- pare PS4 with the two BSRNN-based baseline systems re- leased by the REAL-T challenge, namely BSRNN_TFMAP and BSRNN_EMB. Both models are trained on simulated mixtures and have not been adapted to real conversational recordings. In addition to the development set, we also submit our system to the official REAL-T leaderboard. Note that the challenge organizers updated the perceptual quality metric from DNSMOS OVRL to DNSMOS-P808 [28] for the final leaderboard evaluation. Therefore, we report DNSMOS-P808 scores for the leaderboard results. 6 https://github.com/microsoft/dns-challenge TABLE I RESULTS ON REAL-T CHALLENGE LEADERBOARD (VALIDATION SET). SystemTER↓F1↑SIM↑DNSMOS-P808↑ MERL’s0.6130.8610.5383.371 PS4 (ours)0.6390.8710.5653.128 CARTSE’s0.6510.8570.5443.138 BSRNNEMB0.8290.8290.4172.875 BSRNNTFMAP0.8380.8290.4432.756 Fig. 2 visualizes the per-dataset performance across all four metrics. PS4 consistently outperforms both baselines on every sub-corpus, with the most pronounced gains on DNSMOS OVRL, where the baseline systems score below 2.0 while PS4 exceeds 3.1 on all five datasets. The TER improvements are also consistent, with PS4 achieving lower error rates than both baselines across all sub-corpora. Table I reports the detailed per-dataset numerical results of PS4. AMI achieves the lowest TER of 0.388 and the highest SIM of 0.714, likely because AMI recordings have relatively clean lapel micro- phone channels that benefit enrollment extraction. AISHELL- 4 shows the highest TER of 0.567, which we attribute to the more challenging far-field acoustic conditions and the higher proportion of overlapping speech in that corpus. Across all five datasets, PS4 achieves consistent SIM improvements over the mixture baseline, confirming that the speaker similarity ranking loss effectively guides the TSE model to preserve target speaker identity. Table I shows the challenge leaderboard results on the official validation set. PS4 ranks 2nd overall on the leaderboard with a composite score of 3.25. Notably, PS4 achieves the best F1 of 0.871 and the best SIM of 0.565 among all submitted systems, demonstrating superior speaker extraction accuracy and identity preservation. Although the DNSMOS-P808 score of 3.128 is slightly lower than the top-ranked MERL’s sys- tem, and TER is slightly higher, PS4 decisively outperforms MERL’s on SIM and F1. These results confirm that our four proxy supervision objectives provide effective training signals for TSE in real conversational scenarios, particularly for linguistic accuracy and speaker identity preservation. V. CONCLUSION In this paper, we presented PS4, a proxy-supervised joint training framework for target speaker extraction from real conversational recordings. To address the lack of clean ref- erence speech in real-world data, we constructed REAL- PS4, a large-scale training corpus derived from four public datasets, covering both Chinese and English scenarios. We further proposed a joint optimization strategy that combines four complementary proxy objectives to supervise the TSE model without requiring clean isolated speech. Experiments on the REAL-T benchmark demonstrate that PS4 consistently outperforms the official baselines across all five sub-corpora and all four evaluation metrics. On the official challenge leaderboard, PS4 achieves the best F1 and speaker similarity scores among all submitted systems, ranking 2nd overall. These results validate that proxy supervision provides an effective and practical alternative to conventional signal-level training for TSE in real conversational scenarios. REFERENCES [1] Q. Wang, H. Muckenhirn, K. Wilson, P. Sridhar, Z. Wu, J. Hershey, R. A. Saurous, R. J. Weiss, Y. Jia, and I. L. Moreno, “Voicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,” arXiv preprint arXiv:1810.04826, 2018. [2] S. He, H. Li, and X. Zhang, “Speakerfilter: Deep learning-based target speaker extraction using anchor speech,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, p. 376–380. [3] K. Zmolikova, M. Delcroix, T. Ochiai, K. Kinoshita, J. ˇ Cernock ` y, and D. Yu, “Neural target speech extraction: An overview,” IEEE Signal Processing Magazine, vol. 40, no. 3, p. 8–29, 2023. [4] H. Sato, T. Ochiai, K. Kinoshita, M. Delcroix, T. Nakatani, and S. Araki, “Multimodal attention fusion for target speaker extraction,” in 2021 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2021, p. 778–784. [5] Y. Liu, X. Liu, and J. Yamagishi, “Improving curriculum learning for target speaker extraction with synthetic speakers,” in 2024 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2024, p. 364–370. [6] Y. Luo and N. Mesgarani, “Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,” IEEE/ACM trans- actions on audio, speech, and language processing, vol. 27, no. 8, p. 1256–1266, 2019. [7] Y. Luo, Z. Chen, and T. Yoshioka, “Dual-path rnn: Efficient long sequence modeling for time-domain single-channel speech separation,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, p. 46–50. [8] C. Subakan, M. Ravanelli, S. Cornell, M. Bronzi, and J. Zhong, “Attention is all you need in speech separation,” in ICASSP 2021- 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, p. 21–25. [9] B. Zeng and M. Li, “Usef-tse: Universal speaker embedding free target speaker extraction,” IEEE Transactions on Audio, Speech and Language Processing, 2025. [10] T. Ling, S. He, P. Shen, and Z.-Q. Wang, “Mc-lext: Multi-channel target speaker extraction with onset-prompted speaker conditioning mechanism,” in ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).IEEE, 2026, p. 18 967–18 971. [11] A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: a large-scale speaker identification dataset,” arXiv preprint arXiv:1706.08612, 2017. [12] J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in 2016 IEEE international conference on acoustics, speech and signal process- ing (ICASSP). IEEE, 2016, p. 31–35. [13] J. Cosentino, M. Pariente, S. Cornell, A. Deleforge, and E. Vincent, “Librimix: An open-source dataset for generalizable speech separation,” arXiv preprint arXiv:2005.11262, 2020. [14] J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal et al., “The ami meeting corpus: A pre-announcement,” in International workshop on machine learning for multimodal interaction. Springer, 2005, p. 28– 39. [15] F. Yu, S. Zhang, Y. Fu, L. Xie, S. Zheng, Z. Du, W. Huang, P. Guo, Z. Yan, B. Ma et al., “M2met: The icassp 2022 multi-channel multi- party meeting transcription challenge,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, p. 6167–6171. [16] S. Watanabe, M. Mandel, J. Barker, E. Vincent, A. Arora, X. Chang, S. Khudanpur, V. Manohar, D. Povey, D. Raj et al., “Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,” arXiv preprint arXiv:2004.09249, 2020. [17] Y. Fu, L. Cheng, S. Lv, Y. Jv, Y. Kong, Z. Chen, Y. Hu, L. Xie, J. Wu, H. Bu et al., “Aishell-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,” arXiv preprint arXiv:2104.03603, 2021. [18] M. Van Segbroeck, A. Zaid, K. Kutsenko, C. Huerta, T. Nguyen, X. Luo, B. Hoffmeister, J. Trmal, M. Omologo, and R. Maas, “Dipco–dinner party corpus,” arXiv preprint arXiv:1909.13447, 2019. [19] Shaole Li and Shuai Wang and Jiangyu Han and Ke Zhang and Wupeng Wang and Haizhou Li, “REAL-T: Real Conversational Mixtures for Target Speaker Extraction,” in Interspeech 2025, 2025, p. 1923–1927. [20] J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “Sdr–half-baked or well done?” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).IEEE, 2019, p. 626–630. [21] J. Sohn, N. S. Kim, and W. Sung, “A statistical model-based voice activity detection,” IEEE signal processing letters, vol. 6, no. 1, p. 1–3, 1999. [22] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak super- vision,” arXiv preprint arXiv:2212.04356, 2023. [23] J. Yu, Y. Luo, H. Chen, R. Gu, and C. Weng, “High fidelity speech enhancement with band-split rnn,” arXiv preprint arXiv:2212.00406, 2022. [24] B. Desplanques, J. Thienpondt, and K. Demuynck, “Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,” arXiv preprint arXiv:2005.07143, 2020. [25] A. Mao, M. Mohri, and Y. Zhong, “Cross-entropy loss functions: Theoretical analysis and applications,” in International conference on Machine learning. pmlr, 2023, p. 23 803–23 828. [26] H. Wang, C. Liang, S. Wang, Z. Chen, B. Zhang, X. Xiang, Y. Deng, and Y. Qian, “Wespeaker: A research and production oriented speaker embedding learning toolkit,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).IEEE, 2023, p. 1–5. [27] C. K. Reddy, V. Gopal, and R. Cutler, “Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, p. 6493–6497. [28] B. Naderi and R. Cutler, “An open source implementation of itu-t rec- ommendation p. 808 with validation,” arXiv preprint arXiv:2005.08138, 2020.