Paper deep dive
Speech Enhancement Based on Drifting Models
Liang Xu, Diego Caviedes-Nozal, Bastiaan Kleijn, Longfei Felix Yan, Rasmus Kongsgaard Olsson
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 99%
Last extracted: 6/21/2026, 8:11:35 AM
Summary
The paper introduces DriftSE (Speech Enhancement based on Drifting Models), a novel generative framework that treats speech denoising as a distributional equilibrium problem. Unlike iterative diffusion models, DriftSE achieves one-step inference by evolving a pushforward distribution toward a clean speech distribution using a learned 'Drifting Field'. The framework utilizes a semantic latent space (via SSL encoders like HuBERT, WavLM, or DistilHuBERT) to ensure high-fidelity reconstruction and robust noise suppression. The authors demonstrate two paradigms: direct mapping and a stochastic conditional generative model. Experimental results on the VoiceBank-DEMAND benchmark and DNS Challenge 2020 show that DriftSE outperforms multi-step diffusion baselines in both perceptual quality and generalization, achieving state-of-the-art performance in single-step inference.
Entities (10)
Relation Signals (5)
Liang Xu → affiliatedwith → Victoria University of Wellington
confidence 100% · Liang Xu 1 ... 1 Victoria University of Wellington
DriftSE → evaluatedon → VoiceBank-DEMAND
confidence 100% · Experiments on the VoiceBank-DEMAND benchmark demonstrate that DriftSE achieves high-fidelity enhancement
DriftSE → evaluatedon → DNS Challenge 2020
confidence 100% · We further evaluate the generalization performance on the DNS Challenge 2020 blind test set
DriftSE → uses → Drifting Field
confidence 100% · This evolution is driven by a Drifting Field, a learned correction vector
DriftSE → utilizes → HuBERT
confidence 90% · In particular, we employed HuBERT, WavLM and DistilHuBERT
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem. Rather than relying on iterative sampling, DriftSE natively achieves one-step inference by evolving the pushforward distribution of a mapping function to directly match the clean speech distribution. This evolution is driven by a Drifting Field, a learned correction vector that guides samples toward the high-density regions of the clean distribution, which naturally facilitates training on unpaired data by matching distributions rather than paired samples. We investigate the framework under two formulations: a direct mapping from the noisy observation, and a stochastic conditional generative model from a Gaussian prior. Experiments on the VoiceBank-DEMAND benchmark demonstrate that DriftSE achieves high-fidelity enhancement in a single step, outperforming multi-step diffusion baselines and establishing a new paradigm for speech enhancement.
Tags
Links
- Source: https://arxiv.org/abs/2604.24199v1
- Canonical: https://arxiv.org/abs/2604.24199v1
Trouble viewing inline? Open PDF directly →
Full Text
31,646 characters extracted from source content.
Expand or collapse full text
Submitted to Interspeech 2026. Speech Enhancement Based on Drifting Models Liang Xu 1 , Diego Caviedes-Nozal 2 , Bastiaan Kleijn 1 , Longfei Felix Yan 1 , Rasmus Kongsgaard Olsson 2 1 Victoria University of Wellington, New Zealand 2 GN Audio A/S, Denmark liang.xu,bastiaan.kleijn,felix.yan@vuw.ac.nz,dcnozal,rkolsson@gn.com Abstract We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates de- noising as an equilibrium problem. Rather than relying on it- erative sampling, DriftSE natively achieves one-step inference by evolving the pushforward distribution of a mapping func- tion to directly match the clean speech distribution. This evo- lution is driven by a Drifting Field, a learned correction vec- tor that guides samples toward the high-density regions of the clean distribution, which naturally facilitates training on un- paired data by matching distributions rather than paired sam- ples. We investigate the framework under two formulations: a direct mapping from the noisy observation, and a stochastic conditional generative model from a Gaussian prior. Exper- iments on the VoiceBank-DEMAND benchmark demonstrate that DriftSE achieves high-fidelity enhancement in a single step, outperforming multi-step diffusion baselines and establishing a new paradigm for speech enhancement. Index Terms: speech enhancement, drifting models, diffusion models, consistency models 1. Introduction The field of Speech Enhancement (SE) has evolved significantly over recent decades, progressing from classical statistical sig- nal processing techniques like Wiener filtering [1, 2] to mod- ern deep learning. Discriminative models, such as RNNs [3], LSTMs [4], and complex spectral mapping [5], effectively sup- press noise but often yield spectral oversmoothing and robotic artifacts due to regression-based objectives. Generative Adver- sarial Networks (GANs) [6–8] improve perceptual quality but suffer from training instability and mode collapse. Recently, Score-based Diffusion Models [9] have established state-of-the- art performance by modeling the gradient of the log-density of the clean speech distribution. These models define a forward process that gradually degrades data into noise, and a reverse process for generation. The reverse dynamics can be formulated as either a Stochastic Differential Equation (SDE) [10] or a de- terministic Probability Flow ODE (PF-ODE) sharing the same marginal probability densities. However, their inference is in- herently iterative. Numerically integrating these highly curved reverse-time trajectories requires 10–100 discretization steps, resulting in a high Number of Function Evaluations (NFE) that imposes a critical latency bottleneck for real-time applications. To address the computational inefficiency of diffusion mod- els, recent research broadly falls into two lines of work: trajec- tory compression and trajectory linearization. Compression ap- proaches accelerate sampling by reducing the number of steps. For instance, hybrid approaches [11, 12] combine predictive models with a small number of diffusion refinement steps, while diffusion-GAN hybrids [13] further reduce steps via adversarial training. Similarly, distillation-based one-step generators such as Consistency Models [14–17] enforce self-consistency along the PF-ODE to distill a multi-step sampler into a single-step mapping. In parallel, Flow Matching [18, 19] techniques seek to lin- earize the generative trajectory. Rectified Flow [20] explicitly straightens the transport path to minimize the curvature of the ODE. MeanFlow [21, 22] learns a continuous mean velocity field to model probability paths. However, these methods re- main fundamentally trajectory-based. They rely on continuous transport dynamics that must be discretized at inference, and accurately approximating these paths with only a few steps re- mains challenging. Recently, Drifting Models [23] were proposed as a pow- erful new paradigm that reformulates generation as a distri- butional equilibrium problem. By mapping high-dimensional data into a semantic latent space during training, the framework learns a kernelized drifting field where generated samples are si- multaneously attracted to the true data distribution and repelled by the evolving model distribution. Minimizing this drift di- rectly aligns the generator’s pushforward distribution with the target data. Operating natively in a single step, this latent equi- librium approach has achieved state-of-the-art results in large- scale image generation (FID 1.54 on ImageNet). Our contribution is the introduction of Speech Enhance- ment based on Drifting Models (DriftSE), a novel generative framework for one-step denoising. We redefine enhancement as learning a direct projection onto the clean speech manifold without predefined trajectory constraints. Unlike the original formulation which focused on noise-to-data generation [23], we adapt DriftSE for two distinct enhancement paradigms: a direct mapping that pushes the noisy speech distribution toward the clean speech distribution, and a stochastic conditional genera- tive approach that generates clean speech from a Gaussian noise prior. To construct a perceptually meaningful drifting field, we project the audio into a semantic latent space using a pre-trained Self-Supervised Learning (SSL) encoder. Aligning the gener- ated and clean speech distributions within this latent space en- sures effective noise suppression and high-frequency structural recovery. Extensive experiments on the VoiceBank-DEMAND (VB- DMD) dataset demonstrate that DriftSE achieves competi- tive perceptual quality with single-step inference.Specifi- cally, the direct mapping variant achieves PESQ 3.15 and SI- SDR 16.1 dB, while the conditional variant achieves SCOREQ 4.33. Both achieve high-fidelity enhancement without iterative sampling or predefined trajectories. Evaluation on real-world recordings from the DNS Challenge 2020 blind test set demon- strates state-of-the-art generalization performance. arXiv:2604.24199v1 [cs.SD] 27 Apr 2026 2. Drifting Models We briefly review Drifting Models [23], which formulate gener- ative modeling as the training-time evolution of a pushforward distribution. 2.1. Pushforward and Equilibrium Given a simple source distribution p ε (e.g., standard Gaussian noiseN (0,I)), the drift approach takes a sample ε ∼ p ε with ε ∈ R d , and maps it through a parameterized function f θ : R d → R d in a single step to produce a target variable x∈ R d x = f θ (ε).(1) This defines the pushforward distribution q θ = (f θ ) # p ε , mean- ing that sampling ε ∼ p ε and applying f θ yields samples distributed as q θ . For conditional generation, (1) extends to x = f θ (ε,c), where c is a condition (e.g., a class label or noisy speech). To drive q θ toward the target data distribution p data , the framework introduces a Drifting Field V p,q : R d → R d that acts as a correction vector at each generated sample of q θ , point- ing in the direction that reduces the discrepancy between q θ and p data . Concretely, a drifting target is defined as x target ← x + V p,q (x),(2) and the generator is trained to map ε to x target . As training pro- gresses, the pushforward distribution q θ evolves until it reaches a state of distributional equilibrium where the drift vanishes q θ = p data =⇒ V p,q (x) = 0, ∀x.(3) 2.2. Designing the Drifting Field Inspired by mean-shift theory [24], the total drift V p,q (x) de- composes into two opposing forces V p,q (x) = V + p (x)− V − q (x),(4) where V + p attracts samples toward the data distribution p data , and V − q repels samples away from high-density regions of the current model distribution q θ . Both terms take the form of a kernel-weighted mean shift V + p (x) = 1 Z p (x) E y + ∼p k(x,y + )(y + − x) ,(5) V − q (x) = 1 Z q (x) E y − ∼q k(x,y − )(y − − x) ,(6) with similarity kernel k(x,y) and normalizers Z p (x) = E y + ∼p [k(x,y + )] and Z q (x) = E y − ∼q [k(x,y − )]. In practice, these expectations are approximated with mini- batch averages, drawing positives y + from p data and negatives y − from the current model distribution q θ . By substituting the normalizers and combining the forces into a joint expectation over p and q, the self-referential−x terms elegantly cancel out. This yields the exact unified formulation V p,q (x) = 1 Z p Z q E p,q k(x,y + )k(x,y − )(y + − y − ) . (7) To measure the similarity between a generated sample x and any reference feature y ∈ y + ,y − , an exponential simi- larity kernel with temperature τ is employed k τ (x,y) = exp − ∥x− y∥ 2 τ ,(8) where τ controls the interaction bandwidth. 2.3. Training Objective In practice, the drifting field is computed in a latent space using a pretrained extractor φ(·). To optimize the generator f θ , its outputs x = f θ (ε) are driven along the field V via the objective L drift = E ε φ(x)− sg φ(x) + V φ(x) 2 2 ,(9) where sg(·) is the stop-gradient operator. Regressing toward this fixed target minimizes the magnitude of V, progressively transporting the pushforward distribution q θ toward the target data distribution p data . 3. Speech Enhancement via Latent Drifting We propose DriftSE, which formulates speech enhancement as an equilibrium problem (Fig. 1). By evolving the mapping func- tion’s pushforward distribution to match the clean speech distri- bution, DriftSE achieves native one-step denoising (1 NFE). 3.1. Two Enhancement Paradigms Let y ∈ C F×T denote the complex spectrogram of the noisy speech, with F frequency bins and T time frames, and let x ∈ C F×T be the clean speech target. To produce the en- hanced speech ˆ x, we investigate two distinct formulations for the mapping function f θ : Direct Mapping: Defined as ˆ x = f θ (y + σε), where ε∼N (0,I) and σ controls the noise injection strength. During training, when σ = 0, this acts as a strictly deterministic map- ping ˆ x = f θ (y). When σ > 0, the injected noise smooths the acoustic distribution, aiding the estimation of the drifting field. Conditional Generator: Defined as ˆ x = f θ (ε,y), where the network maps a standard Gaussian noise priorε to the target distribution conditioned on y. 3.2. Speech Latent Encoder Applying the drifting field in the time-frequency domain is sub- optimal, as Euclidean distances on raw spectrograms are domi- nated by high-amplitude harmonics, neglecting the low-energy transients critical for phonetic intelligibility. Following [23], we instead compute drift in a semantic latent space and ap- ply multi-layer latent supervision, which provides a richer and more stable training signal. For speech enhancement, we se- lected self-supervised speech models. In particular, we em- ployed HuBERT, WavLM and DistilHuBERT [25–27], which exhibit a well-documented layer hierarchy: shallow layers cap- ture low-level acoustic structure, while deeper layers encode phonetic and semantic content. We therefore define a frozen self-supervised learning (SSL) encoder Φ : R L → R T ′ ×D that maps a waveform of length L to T ′ frame-level latent features of dimension D. To capture the hierarchical speech structures, the drifting field is computed and aggregated across a selected set of layersS. 3.3. Frame-Wise Latent Drifting and Inference Latent Drifting: As detailed in Fig. 1, we dynamically con- struct a positive set of samplesZ + from clean reference frames Φ(x) and a negative set of samplesZ − from the current batch of generated frames Φ( ˆ x). For any generated frame z i ∈ Z − , the frame-wise Drifting Field V(z i ) is computed by instantiat- ing (7) in the speech latent space, using the multi-temperature kernel k τ from (8). The resulting field combines an attraction Figure 1: Overview of the DriftSE framework (illustrating the Direct Mapping formulation). force pulling z i toward the clean distributionZ + and a repul- sion force pushing it away from the current generated distribu- tionZ − , driving f θ toward equilibrium. Training Objective: To capture hierarchical speech struc- tures, the base drifting loss from (9) is computed and aggre- gated across multiple layers l ∈ S of the latent encoder, where S is the set of selected layers, with each layer equally weighted. Inference: At inference time, the direct mapping approach uses σ = 0 for deterministic denoising, while the conditional generator draws a freshε to generate diverse enhanced outputs. Overview of DriftSE: Fig. 1 illustrates the method. The mapping function f θ processes the noisy speech spectrogram alongside injected Gaussian noiseε, which acts as a distribu- tion smoother, to produce a denoised spectrogram in a single step. After iSTFT, both the enhanced waveform ˆ x and the clean reference x are projected into a frame-wise latent space via a frozen encoder Φ. For frame-wise latent drifting at each train- ing iteration, we dynamically construct a mini-batch positive set Z + sampled from the clean feature frames Φ(x), and a mini- batch negative setZ − sampled from the mapped feature frames Φ( ˆ x). The total Drifting Field V for a mapped frame is com- posed of an attraction force V + p that pulls it toward high-density regions of the empirical target distributionZ + , and a repulsion force V − q that pushes it away from its neighbors in the current model distribution Z − . As training progresses, the mapping function minimizes V, dynamically evolving the pushforward distribution until it matches the clean latent distribution at equi- librium (epoch 100). 4. Experiments In this section, we evaluate DriftSE against state-of-the-art it- erative and one-step baselines, and perform ablation studies to analyze the contribution of each design choice. 4.1. Experimental Setup Datasets: We train on clean speech from the VoiceBank cor- pus [28] and noise recordings from the DEMAND dataset [29]. During training, we employ dynamic mixing where 10,802 clean utterances are mixed on-the-fly with 18 distinct noise types [9]. To ensure robust generalization and prevent over- fitting to specific acoustic conditions, Signal-to-Noise Ratios (SNRs) are sampled randomly from0, 5, 10, 15 dB. For evaluation, we utilize the standard pre-mixed VB-DMD test set (824 utterances) to ensure fair benchmarking. To assess real-world generalization, we further evaluate on the DNS Chal- lenge 2020 blind test set [30], which contains 300 real-world noisy recordings without clean references. Evaluation: Following previous studies [9, 16], we re- port pairwise metrics including PESQ [31], ESTOI [32] and SI-SDR [33]. To assess perceptual quality without clean ref- erence, we also report non-intrusive metrics: SCOREQ [34], DNSMOS [35, 36], and WV-MOS [37]. Implementation Details: We employ the NCSN++V2 ar- chitecture [9] as our backbone, omitting the time embedding. Audio samples are processed at 16 kHz using a Short-Time Fourier Transform (STFT) with a window size of 510, a hop length of 128, and a Hann window, followed by the spectral compression strategy from [9]. For the mapping DriftSE vari- ant, we empirically sample the noise level σ from a truncated log-normal distribution, i.e., logσ ∼ N (−3.0, 1.2), truncated to σ ∈ [0.01, 0.3]. For the conditional generative variant, the STFT spectrogram of the noisy observation y is provided as a conditioning embedding into the generator at each res- olution level. For the SSL latent encoder Φ, we utilize the pre-trained HuBERT-Large, WavLM-Large, and DistilHuBERT checkpoints. The extracted latent frames have a 20ms hop size and 25ms receptive field. We aggregate features from layers S = 6, 12, 24 for WavLM-Large and HuBERT-Large, and layersS = 0, 1, 2 for DistilHuBERT, with all layers equally weighted. We use a multi-temperature exponential kernel ( (8)) with temperatures τ ∈ 0.1, 0.5, 1.0. The model was trained on a single NVIDIA RTX A6000 GPU (48GB VRAM) for 100 epochs. We utilized a batch size of 16 and optimized the network using the AdamW optimizer with a learning rate of 5× 10 −4 and a weight decay of 0.01. 4.2. Results In-domain Evaluation: As shown in Table 1, DriftSE (Dis- tilHuBERT, σ=0) achieves high-fidelity enhancement and out- performs the 30-step SGMSE+ [9] and the one-step Mean- FlowSE [38], reaching PESQ 3.15 and confirming that latent drifting effectively maps noisy observations onto the distribu- tion support of the clean speech. While distillation methods such as ROSE-CD [16] and SBCTM [17] report higher PESQ, they utilize auxiliary losses. When we incorporate the same losses, DriftSE † 1 attains competitive performance. Generalization Evaluation: We evaluate the generaliza- tion capability of DriftSE on the DNS Challenge 2020 blind test set in Table 2. We achieve state-of-the-art WV-MOS 2.65 and SCOREQ 2.97, outperforming other baselines while de- livering highly competitive perceptual scores (DNSMOS SIG, BAK, and OVRL). This confirms that the drifting equilibrium learns a highly generalizable distribution projection. 4.3. Ablations and Analysis Impact of the Latent Encoder: We evaluate the impact of dif- ferent encoders and layer selections on latent drifting. Using only the deepest semantic layer (WavLM, Layer 24) degrades performance, suggesting that highly abstract features miss fine acoustic details. With multi-layer drifting, DistilHuBERT (768- d) is competitive with both HuBERT and WavLM (1024-d), achieving the best SI-SDR while maintaining similar percep- 1 Model jointly trained with auxiliary PESQ and SI-SDR losses. step 999step 1999step 8999step 346999 epoch 1epoch 2epoch 50epoch 100 Figure 2: Evolution of frame-level distributions in the DistilHuBERT semantic space for a fixed test utterance. Each panel displays 2D density contours (PCA projection) derived from all frames across different training epochs. Stars denote the corresponding centroids, which represent the mean of all projected frames. As training progresses, the generated distribution shifts from the noisy distribution toward the clean distribution. tual quality. Therefore, we use DistilHuBERT as the default encoder in subsequent experiments. Conditional Drifting Models (DriftSE ∗ 2 ): As shown in Table 1, DriftSE ∗ successfully reduces generative ambiguity, delivering superior reference-free perceptual metrics (DNS- MOS 3.64, SCOREQ 4.33) while maintaining competitive pair- wise fidelity. This demonstrates that incorporating a stochastic prior enables the generator’s pushforward distribution to better capture the inherent variance of the clean speech distribution, leading to more natural generation. Effect of Noise Injection: For the direct mapping vari- ant, omitting the noise prior (σ = 0) during training enforces a deterministic mapping with higher fidelity (PESQ 3.15, SI- SDR 16.10 dB), whereas injecting Gaussian noise smooths the acoustic distribution to improve reference-free perceptual qual- ity (SCOREQ from 4.08 to 4.15). This distributional smoothing trades marginal waveform precision for more natural genera- tion, providing a promising direction for adapting to narrow or shifted target distributions in future work. Unpaired Training: Here, “unpaired” means the model lacks access to noisy-clean audio pairs during training. For DriftSE (Unpaired, map to DNS), each mini-batch is formed by independently sampling noisy speech from VoiceBank and clean targets from the DNS training set. This still achieves strong reference-free quality (DNSMOS 3.61, SCOREQ 3.92). The expected drop in pairwise fidelity (PESQ 2.00, SI-SDR 6.60 dB) occurs because mini-batch drift estimation is inher- ently less precise than exact paired setting. Nevertheless, the strong non-intrusive scores indicate that the model can still drift its outputs toward the clean speech distribution without access to paired clean targets. For DriftSE (Unpaired, map to VB-Female), noisy inputs contain mixed-gender speech from VoiceBank, while clean tar- gets consist of female speech. This forces the model to learn a female-only target distribution, leading to systematic changes in speaker characteristics. As a result, standard pairwise met- rics are omitted since clean references are no longer aligned with the shifted outputs. The perceptual scores (DNSMOS 3.40, SCOREQ 3.72) indicate successful distribution drift. Distributional Convergence Visualization: We qualita- tively verify the latent drifting mechanism by tracking a fixed test utterance’s evolution in the DistilHuBERT semantic space. Figure 2 reveals a clear transition from a noise distribution to- ward the clean distribution. While the generated audio initially overlaps with the noisy distribution, optimization of the drifting field enables the model to capture the structural characteristics 2 Conditional drifting model. Table 1: Comparison on VB-DMD. NFE: Number of Function Evaluations. Bold indicates the best performing metric within each respective group. MethodNFE PESQ SI-SDR ESTOI DNSMOS SCOREQ MetricGAN+ [39]13.138.500.833.223.82 UNIVERSE++ [40]82.9118.000.853.454.35 SGMSE+ [9]302.9016.900.853.483.98 ROSE-CD [16]13.4917.800.873.494.23 SBCTM [17]13.5612.700.873.554.35 MeanFlowSE [38]12.8119.970.883.584.25 DriftSE (WavLM, L24)12.9012.600.843.363.93 DriftSE (WavLM)13.0314.000.853.544.17 DriftSE (HuBERT)12.9412.500.843.494.14 DriftSE (DistilHuBERT)13.0015.600.853.484.15 DriftSE (DistilHuBERT, σ = 0)13.1516.100.863.474.08 DriftSE ∗ (DistilHuBERT)12.9917.980.863.644.33 DriftSE † (DistilHuBERT)13.4520.600.873.494.11 DriftSE (Unpaired, map to DNS)12.006.600.743.613.92 DriftSE (Unpaired, map to VB-Female)1---3.403.72 Table 2: Real-world recordings evaluation on DNS Challenge 2020 Blind Test Set. Bold indicates the best performing metric within each respective group. MethodNFE WV-MOS SCOREQ SIG BAK OVRL MetricGAN+ [39]11.232.083.283.452.70 UNIVERSE++ [40]81.992.273.453.522.93 SGMSE+ [9]302.342.954.12 3.943.62 ROSE-CD [16]12.372.814.013.803.42 SBCTM [17]12.242.783.833.883.33 MeanFlowSE [38]12.202.793.883.513.21 DriftSE (WavLM)12.622.673.85 3.943.42 DriftSE (HuBERT)12.562.743.923.793.40 DriftSE (DistilHuBERT)12.652.973.783.843.31 DriftSE ∗ (DistilHuBERT)12.452.784.013.683.43 DriftSE † (DistilHuBERT)12.512.864.00 3.823.47 of the clean distribution. The converged contours and centroids demonstrate that DriftSE successfully maps noisy observations to the high-density regions of the clean speech distribution at equilibrium. 5. Conclusion In this paper, we introduced Speech Enhancement based on Drifting Models (DriftSE), a novel paradigm that reformulates denoising as an equilibrium problem to enable native one-step generation. By utilizing a latent drifting field, DriftSE evolves the mapping function’s pushforward distribution to directly match the clean speech distribution during training. We demon- strated that computing this drift within a multi-scale seman- tic latent space provides a robust learning signal that supports high-fidelity reconstruction. Extensive evaluations confirm that DriftSE achieves state-of-the-art perceptual quality and gener- alization, outperforming multi-step diffusion baselines. 6. Generative AI Use Disclosure We acknowledge the ISCA policy stating that generative AI tools cannot serve as co-authors and should only be used for editing or polishing rather than producing significant parts of this paper. Although the proposed method is a novel generative model for speech enhancement, the authors declare that no gen- erative AI tools were used to develop the source code, but AI tools were used to correct text grammar. 7. References [1] J. Meyer and K. U. Simmer, “Multi-channel Speech Enhancement in a Car Environment Using Wiener Filtering and Spectral Sub- traction,” in ICASSP, vol. 2. IEEE, 1997, p. 1167–1170. [2] J. Chua, L. F. Yan, and W. B. Kleijn, “An Effective MVDR Post- Processing Method for Low-Latency Convolutive Blind Source Separation,” in 2024 18th International Workshop on Acoustic Signal Enhancement (IWAENC). IEEE, 2024, p. 130–134. [3] F. Weninger, J. R. Hershey, J. L. Roux, and B. Schuller, “Speech enhancement with LSTM recurrent neural networks and its ap- plication to noise-robust ASR,” in Latent Variable Analysis and Signal Separation (LVA/ICA), vol. 9237, Aug. 2015, p. 91–99. [4] K. Tan and D. Wang, “A Convolutional Recurrent Neural Network for Real-Time Speech Enhancement,” in Proc. Interspeech, 2018, p. 3229–3233. [5] Y. Hu, Y. Liu, S. Lyu, M. Xing, S. Zhang, Y. Fu, J. Fan, and L. Xie, “DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement,” in Proc. Interspeech, 2020, p. 2472–2476. [6] S. Pascual, A. Bonafonte, and J. Serr ` a, “SEGAN: Speech En- hancement Generative Adversarial Network,” in Proc. Inter- speech, 2017, p. 3642–3646. [7] S.-W. Fu, C.-F. Liao, Y. Tsao, and S.-D. Lin, “MetricGAN: Gen- erative Adversarial Networks based Black-box Metric Scores Op- timization for Speech Enhancement,” in International Conference on Machine Learning. PMLR, 2019, p. 2031–2041. [8] J. Su, Z. Jin, and A. Finkelstein, “HiFi-GAN-2: Studio-Quality Speech Enhancement via Generative Adversarial Networks Con- ditioned on Acoustic Features,” in 2021 IEEE Workshop on Appli- cations of Signal Processing to Audio and Acoustics (WASPAA). IEEE, 2021, p. 166–170. [9] J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, “Speech Enhancement and Dereverberation With Diffusion- Based Generative Models,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, p. 2351–2364, 2023. [10] Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Er- mon, and B. Poole, “Score-Based Generative Modeling through Stochastic Differential Equations,” in ICLR, 2021. [11] J.-M. Lemercier, J. Richter, S. Welker, and T. Gerkmann, “StoRM: A Diffusion-Based Stochastic Regeneration Model for Speech Enhancement and Dereverberation,” IEEE/ACM Transac- tions on Audio, Speech, and Language Processing, vol. 31, p. 2724–2737, 2023. [12] T. Trachu, C. Piansaddhayanon, and E. Chuangsuwanich, “Thun- der: Unified Regression-Diffusion Speech Enhancement with a Single Reverse Step using Brownian Bridge,” in Proc. Inter- speech, 2024, p. 1180–1184. [13] S. Han, S. Lee, J. Lee, and K. Lee, “Few-step Adversarial Schr ̈ odinger Bridge for Generative Speech Enhancement,” in Proc. Interspeech, 2025, p. 2380–2384. [14] Y. Song, P. Dhariwal, M. Chen, and I. Sutskever, “Consistency Models,” in Proceedings of the 40th International Conference on Machine Learning, ser. Proceedings of Machine Learning Re- search, vol. 202. PMLR, 23–29 Jul 2023, p. 32 211–32 252. [15] D. Kim, C.-H. Lai, W.-H. Liao, N. Murata, Y. Takida, T. Uesaka, Y. He, Y. Mitsufuji, and S. Ermon, “Consistency Trajectory Mod- els: Learning Probability Flow ODE Trajectory of Diffusion,” in ICLR, 2024. [16] L. Xu, L. F. Yan, and W. B. Kleijn, “Robust One-Step Speech Enhancement via Consistency Distillation,” in Proc. IEEE Work- shop on Applications of Signal Processing to Audio and Acoustics (WASPAA). Tahoe City, CA, USA: IEEE, Oct. 2025. [17] S. Nishigori, K. Saito, N. Murata, M. Hirano, S. Takahashi, and Y. Mitsufuji, “Schr ̈ odinger Bridge Consistency Trajectory Mod- els for Speech Enhancement,” in 2025 IEEE Workshop on Appli- cations of Signal Processing to Audio and Acoustics (WASPAA). IEEE, 2025, p. 1–5. [18] Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow Matching for Generative Modeling,” in ICLR, 2023. [19] Z. Wang, Z. Liu, X. Zhu, Y. Zhu, M. Liu, J. Chen, L. Xiao, C. Weng, and L. Xie, “FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching,” in Proc. Interspeech, 2025. [20] X. Liu, C. Gong, and Q. Liu, “Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow,” in ICLR, 2023. [21] Z. Geng, M. Deng, X. Bai, Z. Kolter, and K. He, “Mean Flows for One-step Generative Modeling,” in NeurIPS, 2025. [22] Z. Geng, Y. Lu, Z. Wu, E. Shechtman, J. Z. Kolter, and K. He, “Improved Mean Flows: On the Challenges of Fastforward Gen- erative Models,” arXiv preprint arXiv:2512.02012, 2025. [23] M. Deng, H. Li, T. Li, Y. Du, and K. He, “Generative Modeling via Drifting,” arXiv preprint arXiv:2602.04770, 2026. [24] Y. Cheng, “Mean shift, mode seeking, and clustering,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 17, no. 8, p. 790–799, 1995. [25] W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhut- dinov, and A.-r. Mohamed, “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” IEEE/ACM Transactions on Audio, Speech, and Language Pro- cessing, vol. 29, p. 3451–3460, 2021. [26] S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al., “WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, p. 1505–1518, 2022. [27] H.-J. Chang, S.-w. Yang, and H.-y. Lee, “DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden- unit BERT,” in ICASSP. IEEE, 2022, p. 7087–7091. [28] C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “In- vestigating RNN-based speech enhancement methods for noise- robust Text-to-Speech,” in 9th ISCA Workshop on Speech Synthe- sis Workshop (SSW 9), 2016, p. 146–152. [29] J. Thiemann, N. Ito, and E. Vincent, “The Diverse Environments Multi-channel Acoustic Noise Database (DEMAND): A database of multichannel environmental noise recordings,” in Proceedings of Meetings on Acoustics, vol. 19, no. 1. AIP Publishing, 2013. [30] C. K. A. Reddy, V. Gopal, R. Cutler, E. Beyrami, R. Cheng, H. Dubey, S. Matusevych, R. Aichner, A. Aazami, S. Braun, P. Rana, S. Srinivasan, and J. Gehrke, “The Interspeech 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results,” in Proc. Interspeech, 2020, p. 2492–2496. [31] A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Per- ceptual Evaluation of Speech Quality (PESQ): A New Method for Speech Quality Assessment of Telephone Networks and Codecs,” in ICASSP, vol. 2. IEEE, 2001, p. 749–752. [32] J. Jensen and C. H. Taal, “An Algorithm for Predicting the In- telligibility of Speech Masked by Modulated Noise Maskers,” IEEE/ACM Transactions on Audio, Speech, and Language Pro- cessing, vol. 24, no. 11, p. 2009–2022, 2016. [33] J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “Sdr – half-baked or well done?” in ICASSP, 2019, p. 626–630. [34] A. Ragano, J. Skoglund, and A. Hines, “SCOREQ: Speech Qual- ity Assessment with Contrastive Regression,” in NeurIPS, vol. 37, 2024, p. 105 702–105 729. [35] C. K. A. Reddy, V. Gopal, and R. Cutler, “DNSMOS: A Non- Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,” in ICASSP. IEEE, August 2021. [36] C. K. Reddy, V. Gopal, and R. Cutler, “DNSMOS P.835: A Non- Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,” in ICASSP. IEEE, 2022, p. 886–890. [37] P. Andreev, A. Alanov, O. Ivanov, and D. Vetrov, “HiFi++: a Uni- fied Framework for Bandwidth Extension and Speech Enhance- ment,” in ICASSP. IEEE, 2023, p. 1–5. [38] D. Li, S. Lu, H. Pan, Z. Zhan, Q. Hong, and L. Li, “Mean- FlowSE: one-step generative speech enhancement via conditional mean flow,” in ICASSP, 2026. [39] S.-W. Fu, C. Yu, Y. Tsao, X. Lu, and H. Kawahara, “MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement,” in Proc. Interspeech, 2021, p. 201–205. [40] R. Scheibler, Y. Fujita, Y. Shirahata, and T. Komatsu, “UNI- VERSE++: Universal Score-based Speech Enhancement with High Content Preservation,” in Proc. Interspeech, 2024.