Paper deep dive
DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation
Wei Zhou, Wanyi Ning, Yinshang Guo, Qianxiao Fang, Haitao Qian, Yingpeng Li
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions. Existing approaches usually fuse visual features directly into the separation network, making them vulnerable to degraded visual signals. In this paper, we present DAVE, a decoupled audio-visual enhancement framework for real-world speech separation. Firstly, to address the data scarcity issue, we construct DAVE-Corpus, a large-scale training corpus with 219,411 mixtures generated from public meeting corpora through combinatorial acoustic augmentation. Then, we introduce a progressive multi-objective optimization strategy to jointly improve speech separation, intelligibility, speaker identity preservation, and perceptual quality. We further develop a certified selective enhancement chain that applies scene routing, GAN-based denoising, and loudness normalization only within the no-reference partition, guaranteeing non-degradation of reference-based metrics. Experimental results on the Real-World Audio-Visual Speech Enhancement Challenge demonstrate the robustness of DAVE under both real-world mixed scenarios and visual degradation conditions.
Tags
Links
- Source: https://arxiv.org/abs/2608.09288v1
- Canonical: https://arxiv.org/abs/2608.09288v1
Trouble viewing inline? Open PDF directly →
Full Text
29,684 characters extracted from source content.
Expand or collapse full text
DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Wei Zhou 1 , Wanyi Ning 1,2,∗ , Yinshang Guo 3 , Qianxiao Fang 1 , Haitao Qian 1 , Yingpeng Li 1 1 Yijiahe AI 2 Tianjin University 3 Nanjing University zhouwei@yijiahe.com, ningwanyi@126.com Abstract Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic condi- tions. Existing approaches usually fuse visual features directly into the separation network, making them vulnerable to de- graded visual signals. In this paper, we present DAVE, a de- coupled audio-visual enhancement framework for real-world speech separation.Firstly, to address the data scarcity is- sue, we construct DAVE-Corpus, a large-scale training corpus with 219,411 mixtures generated from public meeting corpora through combinatorial acoustic augmentation. Then, we in- troduce a progressive multi-objective optimization strategy to jointly improve speech separation, intelligibility, speaker iden- tity preservation, and perceptual quality. We further develop a certified selective enhancement chain that applies scene rout- ing, GAN-based denoising, and loudness normalization only within the no-reference partition, guaranteeing non-degradation of reference-based metrics. Experimental results on the Real- World Audio-Visual Speech Enhancement Challenge 1 demon- strate the robustness of DAVE under both real-world mixed sce- narios and visual degradation conditions. Index Terms: audio-visual speech enhancement, speech separation, target speaker extraction, multi-objective optimiza- tion 1. Introduction Audio-visual speech enhancement (AVSE) aims to recover the speech signal of a target speaker from noisy and interfering speech by exploiting complementary acoustic and visual in- formation. Compared with audio-only speech separation ap- proaches [1, 2, 3, 4], AVSE benefits from visual cues such as facial movements and speaker appearance, which provide valuable information for identifying and extracting the target speaker [5, 6]. Therefore, AVSE has been widely studied for applications including assistive listening, human-computer in- teraction, and multi-speaker communication systems. However, deploying AVSE systems in real-world environments remains challenging. Unlike conventional benchmarks based on clean videos and artificially mixed speech, real-world scenarios often involve naturally occurring speech overlap, reverberation, envi- ronmental noise, and degraded visual observations. Real-world challenges such as the CHiME-5 task [7] highlight the difficulty of multi-party meeting recognition under noisy, far-field condi- tions. These factors create significant challenges for building robust AVSE models. Existing AVSE approaches mainly focus on designing ef- fective audio-visual fusion mechanisms. Typically, audio and ∗ Corresponding author. 1 https://real-world-avse.github.io/ visual representations are extracted separately and then fused within the separation network, allowing visual information to directly guide speech reconstruction [8]. Representative separa- tion backbones include Conv-TasNet [1], DPRNN [9], and Sep- Former [10], which have achieved strong performance on syn- thetic benchmarks. Although such approaches achieve promis- ing performance under controlled conditions, they heavily rely on the quality and reliability of visual inputs. In real-world sce- narios, visual signals may be affected by occlusion, low reso- lution, missing frames, blur, or far-field recording conditions. When degraded visual features are directly incorporated into the separation process, inaccurate visual information may in- terfere with speech reconstruction and reduce system robust- ness. Moreover, the lack of large-scale training data with real- istic acoustic conditions further limits the generalization ability of existing AVSE systems, as models trained on simulated mix- tures often suffer from a considerable domain gap when applied to real-world recordings. In this paper, we propose DAVE, a decoupled audio-visual enhancement framework for real-world speech separation. Un- like conventional fusion-based approaches, DAVE separates au- dio reconstruction from visual speaker attribution, where the au- dio branch focuses on robust speech separation and the visual branch provides speaker-related information for target identifi- cation. Furthermore, we construct DAVE-Corpus, containing 219,411 realistic mixtures generated from public meeting cor- pora [11, 12, 13] through combinatorial acoustic augmentation with room impulse response simulation [14]. Based on DAVE- Corpus, we develop a progressive multi-objective optimization strategy and a metric-preserving selective enhancement chain to improve separation quality, intelligibility, speaker identity preservation, and perceptual quality. The main contributions of this paper are summarized as fol- lows: (1) we construct DAVE-Corpus, a large-scale AVSE train- ing corpus with 219,411 realistic mixtures, providing diverse acoustic conditions for robust model training; (2) we propose DAVE, a decoupled audio-visual enhancement framework with progressive multi-objective optimization and metric-preserving selective enhancement, improving robustness against unreliable visual inputs; and (3) we conduct extensive experiments on the Real-World Audio-Visual Speech Enhancement Challenge un- der realistic mixed scenarios and degraded visual conditions, demonstrating the effectiveness of DAVE in challenging real- world environments. 2. Data Construction To train the audio separation backbone of DAVE under realistic acoustic conditions, we construct a large-scale audio training corpus, termed DAVE-Corpus, in three stages: synthetic mix- ture generation, official protocol remix, and EchoSet replay. Synthetic Mixture Generation. We collect speech seg- arXiv:2608.09288v1 [cs.SD] 10 Aug 2026 Table 1: Synthetic-real performance discrepancy analysis using the TIGER probe model. The gap is measured as the absolute difference between synthetic and real-domain performance. Training DataSyntheticRealAbsolute Gap Additive-only9.3910.701.31 Additive fine-tuned14.0510.493.56 DAVE-Corpus11.1510.700.45 ments from three public meeting corpora: AliMeeting [11], MISP [12], and AISHELL-4 [13]. Since the original recordings contain different levels of noise and interference, we first per- form quality-based screening using frame-level energy statis- tics [15]. Segments are retained if their estimated SNR is above 25 dB, duration is between 2.0 and 15.0 seconds, and silence ratio is below 35%. After filtering, 4,431 speech segments are selected, including 1,469 from AliMeeting, 2,926 from MISP, and 36 from AISHELL-4. A key challenge in real-world AVSE is the acoustic mis- match between conventional additive mixtures and real record- ings. To motivate our augmentation strategy, we first ana- lyze this mismatch using a pre-trained TIGER [16] separation model. As shown in Table 1, models trained with additive- only mixtures exhibit a substantial discrepancy between syn- thetic and real-domain performance. Further fine-tuning on ad- ditive mixtures improves synthetic performance but leads to a larger real-domain gap, suggesting that models may overfit to simplified acoustic conditions. Based on this observation, we design a combinatorial augmentation pipeline to generate di- verse meeting-like mixtures. For each mixture, two speech segments are randomly sampled while avoiding highly similar speakers. Specifically, speaker embeddings are extracted using a pretrained speaker encoder, and a pair is rejected if the cosine similarity exceeds 0.5. The mixture duration is randomly sam- pled between 4.0 and 8.0 seconds, and each source is truncated or repeated accordingly. The second source is randomly scaled with a relative gain sampled from [−3, 3] dB to simulate differ- ent speaker distances. In addition, we use a frozen Fun-ASR- Nano-2512 model [17] to generate reference transcriptions for all training mixtures, providing the text labels required by the CER loss during training. To model realistic acoustic environments, 50% of the mix- tures are augmented with simulated reverberation using py- roomacoustics [18, 14]. Room configurations are randomly sampled with dimensions of 3–8 m × 3–7 m × 2.5–3.5 m and RT60 values between 0.2 and 0.4 seconds. Each source is convolved with its corresponding room impulse response be- fore mixing. The remaining mixtures use additive mixing only. Environmental noise from MUSAN [19] is further added with randomly selected SNR values between 12 and 30 dB. The final synthetic subset contains 199,998 mixtures. Official Protocol Remix. The challenge development set provides a remix partition where mixtures are created by di- rectly adding two single-speaker waveforms following the offi- cial protocol. Using this protocol, we re-synthesize 8,000 train- ing mixtures by randomly pairing de-duplicated single-speaker segments from the development set, while excluding 150 held- out sessions reserved for validation. This anchors the training distribution to the evaluation protocol. EchoSet Replay. The TIGER [16] separation model was pre-trained on the EchoSet corpus, which contains synthetic mixtures with realistic reverberation. To prevent catastrophic Mixed Audio x(t) Face Tracks v 1 , v 2 TIGER-M Audio Separation (2.56M) Two Streams ˆs 1 ,ˆs 2 Voiceprint WeSpeaker w= 3.0 StableSyncNet w= 1.5 SyncNet w= 1.0 Keypoints w= 0.7 Weighted Fusion (normalize→one-shot vote) Speaker Attribution ˆs A ,ˆs B Scene Split (routing) GAN Denoise MossFormerGAN Loudness Norm (target×6) Enhanced Speech ˆs ⋆ A ,ˆs ⋆ B Audio-only backbone Visual decision branch (weighted fusion) Certified selective enhancement Figure 1: Overview of the DAVE framework.The audio- only separation backbone reconstructs two anonymous speech streams, while a visual decision branch assigns speaker identi- ties by weighted fusion of four evidence votes: voiceprint as the primary anchor, StableSyncNet, SyncNet, and lip-keypoint mo- tion. The certified selective enhancement chain applies scene routing, GAN-based denoising, and loudness normalization be- fore output. forgetting [20] of the pre-training domain during fine-tuning on the new DAVE-Corpus, we include 11,413 replay pairs from the EchoSet corpus as an auxiliary subset. These replay pairs are mixed into the training data alongside the synthetic and of- ficial protocol mixtures, preserving the model’s ability to gen- eralize across acoustic domains. The final DAVE-Corpus totals 219,411 training mixtures across three subsets: 199,998 syn- thetic pairs bridging the synthetic-real domain gap, 8,000 offi- cial protocol remixes anchoring the training distribution to the evaluation protocol, and 11,413 EchoSet replay pairs preventing catastrophic forgetting of the pre-training domain. 3. Methodology DAVE decouples audio reconstruction from visual speaker attribution, reserving visual information for low-bandwidth decision-making tasks where it is most reliable. As illustrated in Figure 1, the framework comprises three components. First, an audio-only separation module is trained on the DAVE-Corpus with a progressive multi-objective strategy that jointly opti- mizes separation quality, speech recognition accuracy, speaker identity preservation, and perceptual quality. Second, a vi- sual speaker attribution module assigns speaker identities to the output streams by weighted fusion of audio-visual evi- dence, without influencing the separation process itself. Third, a certified selective enhancement module applies scene rout- ing, GAN-based denoising, and loudness normalization only within the no-reference partition, guaranteeing non-degradation of reference-based metrics. 3.1. Audio Separation We adopt TIGER [16], a lightweight time-frequency domain separation network, as the audio backbone. The original ar- chitecture with 0.82M parameters exhibits capacity saturation, where adding more training data no longer improves separation performance. We therefore scale the model to 2.56M param- eters by increasing the encoder channels from 256 to 512, the feature channels from 128 to 256, and the number of repeating blocks from 8 to 12, while keeping the analysis window of 640 samples and hop size of 160 samples at 16 kHz unchanged. The resulting model, termed TIGER-M, is trained from scratch on the DAVE-Corpus. We train TIGER-M with a multi-objective optimization strategy that jointly optimizes separation quality, speech recognition accuracy, speaker identity preservation, and perceptual quality. Permutation-invariant SI-SDR loss. The primary separa- tion loss is the negative scale-invariant signal-to-distortion ra- tio [21] under permutation-invariant training [2]: L PIT = min π∈Π X i −SI-SDR(ˆs π(i) , s i ),(1) where Π denotes the set of permutations over the two output streams and s i are the reference signals. CER loss. To improve speech recognition accuracy, we minimize the teacher-forcing cross-entropy of a frozen Fun- ASR-Nano-2512 ASR model [17] on the separated outputs: L cer = 1 2 X i CE ASR(ˆs i ),y i ,(2) where y i is the reference transcription. The gradient from text- level loss is back-propagated to the waveform through a differ- entiable log-mel filterbank front-end. The same ASR model is used for both self-labeling the training data and official CER evaluation, ensuring that the loss only penalizes degradations that are measurable by the evaluation metric. Speaker fidelity loss. To preserve speaker identity, we min- imize the cosine distance between the speaker embeddings of the estimated and reference signals: L spk = 1− 1 2 X i cos e(ˆs i ),e(s i ) ,(3) where e(·) is the speaker embedding extracted by a frozen WeS- peaker ResNet34 [22] model, which is the official speaker sim- ilarity evaluation model. The embedding function is made dif- ferentiable by bypassing its default inference-only wrapper. Perceptual losses. We further incorporate differentiable implementations of STOI [23] and PESQ [24], each validated against the official metric with Pearson correlations of 1.0000 and 0.9249, respectively. A differentiable UTMOS [25] critic is implemented by reusing the official UTMOSv2 weights in a trainable-compatible forward pass, with gradient clipping and hybrid-precision safeguards. These losses are computed only on samples with clean reference signals. 3.2. Visual Speaker Attribution The audio separation module outputs two anonymous streams without speaker identities. A visual speaker attribution module assigns each stream to its corresponding speaker by weighted fusion of multiple audio-visual evidence votes, without modi- fying the separation outputs. Instead of relying on a single detector, we combine four complementary evidence sources into a single weighted vote. The first is the voiceprint similarity between each separated stream and the surrounding speaker context, extracted with We- Speaker ResNet34 [22] trained on CN-Celeb [26], the same model family used for official speaker similarity evaluation; this evidence carries the highest weight and serves as the primary anchor of the attribution. The second is the audio-visual match- ing score from StableSyncNet [27], a stable lip-sync discrimina- tor that provides complementary audio-visual correspondence cues. The third is the audio-visual synchrony score from a pre-trained SyncNet [28], fine-tuned following the contrastive cross-modal paradigm of Perfect Match [29]. The fourth is a synchrony metric constructed from lip-keypoint motion en- ergy extracted by FAN [30], which is robust to degraded pixel content. Each vote is normalized and combined with its cor- responding weight, namely 3.0 for voiceprint, 1.5 for audio- visual matching, 1.0 for SyncNet, and 0.7 for keypoints, to pro- duce a single one-shot attribution decision in a single pass. 3.3. Certified Selective Enhancement The separated outputs after speaker attribution still retain resid- ual noise from the original recording. As shown in Figure 1, we apply a final enhancement chain in the certified selective en- hancement module, which is deliberately restricted to the no- reference partition so that reference-based metrics are never modified by construction. Scene Routing. The official evaluation only measures the reference-based metrics SI-SDR, PESQ, and STOI on the syn- thetic remix subset, whereas the perception-oriented metrics are evaluated on all partitions. This motivates restricting any en- hancement to the no-reference partition. We train a lightweight acoustic scene classifier that distinguishes remix from real- recording samples using spectral statistics, achieving 98.13% accuracy on the development set.Only samples routed to the no-reference class enter the subsequent enhancement steps, while samples with ground-truth references are passed through unchanged. This yields a structural guarantee that the reference- based metrics cannot degrade. Layered GAN denoising.The residual environmental noise in the no-reference outputs degrades the perception- oriented metrics. However, applying denoising indiscriminately risks distorting speech content and lowering other measured metrics. We adopt a layered GAN-based denoiser [31] built on MossFormerGAN [32], which is applied only to the routed no- reference samples and denoises the residual interferer without altering the target speech content, improving perceived quality while keeping the other scored metrics stable. Loudness Normalization. Finally, we apply a loudness normalization stage that scales each output by a factor of 6× relative to the original amplitude, following the ITU-R BS.1770 loudness measurement [33]. The gain is taken as the smaller of the target multiple and the peak-protection bound 0.99/peak to prevent clipping, and it is verifiably lossless with respect to the speaker-similarity metric.This final stage improves the DNSMOS-OVRL [34] score and stabilizes loudness across samples without affecting the reference-based metrics, which are confined to the untouched remix partition. 4. Experiments The TIGER-M model is trained from scratch on the DAVE- Corpus using 8×A800 GPUs with native PyTorch DDP. Each GPU processes a batch of 2 three-second segments. We use the Adam optimizer [35] with a peak learning rate of 1 × 10 −3 and a linear warm-up of 1,000 steps, followed by ReduceL- ROnPlateau scheduling with patience of 3. The model is eval- uated every 4,000 steps on a held-out validation set of 149 ses- sions from the development set, using permutation-invariant SI- SDR [21] as the selection criterion. The evaluation protocol fol- lows the official Real-World AVSE Challenge, reporting seven metrics: SI-SDR [21], PESQ [24], STOI [23], UTMOS [25], DNSMOS OVRL [34], CER, and speaker similarity [22]. The Table 2: Stepwise ablation study on the official development set. ConfigurationSI-SDR↑PESQ↑STOI↑UTMOS↑DNS-OVRL↑CER↓SPK↑ Baseline (PIT SI-SDR)9.842.500.7991.9991.62421.9%0.663 + CER loss10.312.630.8101.9991.63119.3%0.679 + Speaker fidelity10.312.630.8101.9991.63119.3%0.697 + Perceptual losses10.232.720.8172.772.0217.1%0.726 + Certified enhancement10.232.720.8172.812.1217.1%0.726 Table 3: Official challenge leaderboard results on the test set for the top 5 teams per track. Track 1: Real-World Mixed; Track 2: Visual Degradation. Rank is the mean rank across the seven metrics; lower is better. CER is reported as a fraction; lower is better. TrackSystemSI-SDR↑PESQ↑STOI↑UTMOS↑DNS-OVRL↑CER↓SPK↑Rank↓ Track 1 Official baseline −5.931.1370.3040.8121.4471.0250.328— audioman10.702.9330.8412.0741.8320.1230.7492.86 AITD12.72 3.000 0.8522.1001.6970.145 0.7692.86 twilight9.162.7750.8252.1451.9440.1350.7423.57 DAVE (ours)10.232.7170.8172.7652.0220.1710.7264.00 SUSTechAILab10.422.6510.8231.9981.7700.1640.7375.43 Track 2 Official baseline −1.691.3040.5021.1001.2071.0520.396— AITD12.26 2.939 0.8442.1291.6700.153 0.7642.43 audioman10.222.8880.8302.0941.7600.1480.7442.86 DAVE (ours)8.932.6150.7892.7662.0120.2200.7074.14 twilight7.052.5610.7762.1701.8960.2200.7125.00 SUSTechAILab9.542.5540.8052.0191.7330.2110.7265.29 final ranking is determined by the mean rank across all seven metrics. Ablation Study. We conduct stepwise ablation experi- ments to measure the contribution of each component. Table 2 summarizes the incremental impact of the training objectives and the certified selective enhancement on the official develop- ment set. The baseline TIGER-M trained with only the PIT SI- SDR loss achieves an SI-SDR of 9.84 dB and a CER of 21.9% on the development set. Adding the CER loss improves the CER from 21.9% to 19.3%, and also yields a side benefit of +0.47 dB in SI-SDR and +0.016 in speaker similarity, demonstrating that recognition-aware training provides complementary regulariza- tion to the separation objective. Introducing the speaker fi- delity loss further improves the speaker similarity from 0.679 to 0.697, confirming the effectiveness of training-side speaker identity preservation without degrading other metrics. The per- ceptual losses, including differentiable STOI, PESQ, and UT- MOS, bring substantial improvements to the perceptual met- rics: PESQ increases by 0.09, UTMOS from 1.999 to 2.77, and OVRL from 1.631 to 2.02, while CER drops further to 17.1%. The certified selective enhancement module, by design, only modifies outputs in the no-reference partition where reference- based metrics are not applicable. Therefore, its contribution is reflected exclusively in the reference-free metrics. On the de- velopment set, the scene-routed GAN denoising and loudness normalization improve DNSMOS OVRL by 0.10 and UTMOS by 0.04, while the remaining five metrics are unchanged by con- struction. Official Challenge Results. We evaluate the complete DAVE system on the Real-World AVSE Challenge test set. Ta- ble 3 summarizes the results across both tracks. On Track 1, the real-world mixed track, DAVE substantially outperforms the official baseline across all seven metrics, as detailed in Ta- ble 3. In particular, it raises SI-SDR by over 16 dB, reduces CER from 1.025 to 0.171, and more than doubles speaker sim- ilarity up to 0.726. On Track 2, the visual degradation track, DAVE maintains strong performance despite degraded visual conditions such as occlusion, blur, and missing face tracks. Re- lying on the audio-only separation backbone and the weighted- fusion attribution that is robust to degraded visual evidence, it achieves an SI-SDR of 8.93 dB and a CER of 0.220, confirm- ing the robustness of the decoupled design. The decoupled de- sign demonstrates strong robustness: the separation backbone operates without visual inputs, making it immune to degraded visual conditions. The visual speaker attribution and certified selective enhancement modules provide complementary gains without introducing modality-specific failure modes. 5. Conclusion We presented DAVE, a decoupled audio-visual enhancement framework for real-world speech separation that reserves visual information for reliable decision-making instead of direct fea- ture fusion. To address data scarcity, we built the DAVE-Corpus with 219,411 mixtures via combinatorial acoustic augmentation and room impulse response simulation, and introduced a pro- gressive multi-objective optimization that jointly improves sep- aration, intelligibility, speaker identity, and perceptual quality. A visual speaker attribution module and a certified selective en- hancement chain add robustness without risking metric degra- dation. Experiments on the Real-World AVSE Challenge con- firm that DAVE performs strongly under both real-world mixed and degraded visual conditions, effectively mitigating unreli- able visual inputs and offering a practical architecture for real- world audio-visual speech enhancement. Future work will ex- plore more adaptive decoupled architectures for complex real- world audio-visual environments. 6. References [1] Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time- frequency magnitude masking for speech separation,” IEEE/ACM Trans. Audio, Speech, Lang. Process., vol. 27, no. 8, p. 1256– 1266, 2019. [2] D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,” in Proc. ICASSP, 2017. [3] J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clus- tering: Discriminative embeddings for segmentation and separa- tion,” in Proc. ICASSP, 2016. [4] W. Ning, W. Zhou, Y. Li, Y. Guo, H. Qian, and Y. Cheng, “Ps4: Proxy-supervised joint training for real target speaker extraction,” 2026. [Online]. Available: https://arxiv.org/abs/2607.08111 [5] A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein, “Looking to listen at the cock- tail party: A speaker-independent audio-visual model for speech separation,” ACM Trans. Graph., vol. 37, no. 4, 2018. [6] T. Afouras, J. S. Chung, and A. Zisserman, “The conversation: Deep audio-visual speech enhancement,” in Proc. Interspeech, 2018. [7] J. Barker et al., “The fifth CHiME speech separation and recogni- tion challenge: Dataset, task and baselines,” in Proc. Interspeech, 2018. [8] J. Wu et al., “Time domain audio visual speech separation,” in Proc. ASRU, 2019. [9] Y. Luo, Z. Chen, and T. Yoshioka, “Dual-path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,” in Proc. ICASSP, 2020. [10] C. Subakan, M. Ravanelli, S. Cornell, M. Bronzi, and J. Zhong, “Attention is all you need in speech separation,” in Proc. ICASSP, 2021. [11] Z. Du et al., “The AliMeeting corpus: A multi-modal meeting corpus,” in Proc. ICASSP, 2022. [12] G. Wang et al., “MISP: A multi-modal interactive speech process- ing system,” in Proc. Interspeech, 2021. [13] Y. Fu et al., “AISHELL-4: An open source dataset for speech separation, diarization and recognition,” in Proc. ICASSP, 2021. [14] R. Scheibler, E. Bezzam, and I. Dokmani ́ c, “Pyroomacoustics: A Python package for audio room simulation and array processing algorithms,” in Proc. ICASSP, 2018. [15] L. R. Rabiner, “A tutorial on hidden markov models and selected applications in speech recognition,” Proceedings of the IEEE, 1989. [16] L. Xu, C. Li, and R. Hu, “Tiger: Time-frequency interleaved gain extraction and reconstruction for efficient speech separation,” in Proc. ICASSP, 2025. [17] Z. Gao et al., “FunASR: A fundamental end-to-end speech recog- nition toolkit,” in Proc. Interspeech, 2023. [18] E. A. Habets, “Room impulse response generator,” Technische Universiteit Eindhoven, Tech. Rep, vol. 2, no. 2.4, p. 1, 2006. [19] D. Snyder, G. Chen, and D. Povey, “MUSAN: A music, speech, and noise corpus,” in arXiv:1510.08484, 2015. [20] J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Des- jardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska- Barwinska et al., “Overcoming catastrophic forgetting in neural networks,” Proceedings of the national academy of sciences, vol. 114, no. 13, p. 3521–3526, 2017. [21] J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR— half-baked or well done?” in Proc. ICASSP, 2019. [22] H. Wang, C. Liang, S. Wang et al., “WeSpeaker: A research and production oriented speaker embedding learning toolkit,” in Proc. ICASSP, 2023. [23] C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An al- gorithm for intelligibility prediction of time-frequency weighted noisy speech,” IEEE Trans. Audio, Speech, Lang. Process., vol. 19, no. 7, p. 2125–2136, 2011. [24] A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Per- ceptual evaluation of speech quality (PESQ)—a new method for speech quality assessment of telephone networks and codecs,” in Proc. ICASSP, 2001. [25] T. Saeki et al., “UTMOS: UTokyo-SaruLab system for VoiceMOS challenge 2022,” in Proc. Interspeech, 2022. [26] Y. Fan, J. W. Kang, L. T. Li, K. C. Li, H. L. Chen, S. T. Cheng, P. Y. Zhang, Z. Y. Zhou, Y. Q. Cai, and D. Wang, “CN-Celeb: A challenging Chinese speaker recognition dataset,” in Proc. ICASSP, 2020, p. 7604–7608. [27] C. Li, C. Zhang, W. Xu, J. Xie, W. Feng, B. Peng, and W. Xing, “LatentSync: Taming audio-conditioned latent diffusion models for lip sync with SyncNet supervision,” arXiv preprint arXiv:2412.09262, 2024. [28] J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in ACCV Workshops, 2016. [29] S.-W. Chung, J. S. Chung, and H.-G. Kang, “Perfect match: Improved cross-modal embeddings for audio-visual synchronisa- tion,” in Proc. ICASSP, 2019, p. 3965–3969. [30] A. Bulat and G. Tzimiropoulos, “How far are we from solving the 2D & 3D face alignment problem? (and a dataset of 230,000 3D facial landmarks),” in Proc. ICCV, 2017, p. 1021–1030. [31] S. Pascual, A. Bonafonte, and J. Serr ` a, “SEGAN: Speech en- hancement generative adversarial network,” in Proc. Interspeech, 2017. [32] S. Zhao et al., “MossFormer2: Combining transformer and RNN-free recurrent network for enhanced time-domain monaural speech separation,” in Proc. ICASSP, 2024. [33] International Telecommunication Union, “Recommendation ITU- R BS.1770-4: Algorithms to measure audio programme loudness and true-peak audio level,” ITU, 2015. [34] C. K. Reddy, V. Gopal, and R. Cutler, “DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise sup- pressors,” in Proc. ICASSP, 2021. [35] D. P. Kingma and J. Ba, “Adam: A method for stochastic opti- mization,” in Proc. International Conference on Learning Repre- sentations (ICLR), 2015.