Paper deep dive
Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones
Oshan A. B. Yalegama, Wageesha N. Manamperi
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and sources, shows promising performance when applied to speech enhancement in noisy environments. Estimating the ReTM of sound sources by exploiting the covariance matrices of multichannel recordings is highly beneficial for practical applications and, to date, remains the only proposed approach. This paper investigates deep learning-based ReTM estimation. We propose three novel supervised learning frameworks using time and short-time frequency transform domain convolutional networks, and a Long Short-Term Memory-based recurrent neural network. Experimental results demonstrate that the proposed models achieve more accurate estimation of the ReTM using five objective metrics compared to the covariance-based method. We also show the effectiveness of the proposed frameworks for speech enhancement, achieving performance on par with the baseline method.
Tags
Links
- Source: https://arxiv.org/abs/2608.11627v1
- Canonical: https://arxiv.org/abs/2608.11627v1
Trouble viewing inline? Open PDF directly →
Full Text
32,092 characters extracted from source content.
Expand or collapse full text
Yalegama Manamperi Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones Oshan A. B Wageesha N Abstract The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and sources, shows promising performance when applied to speech enhancement in noisy environments. Estimating the ReTM of sound sources by exploiting the covariance matrices of multichannel recordings is highly beneficial for practical applications and, to date, remains the only proposed approach. This paper investigates deep learning-based ReTM estimation. We propose three novel supervised learning frameworks using time and short-time frequency transform domain convolutional networks, and a Long Short-Term Memory-based recurrent neural network. Experimental results demonstrate that the proposed models achieve more accurate estimation of the ReTM using five objective metrics compared to the covariance-based method. We also show the effectiveness of the proposed frameworks for speech enhancement, achieving performance on par with the baseline method. keywordsRelative transfer matrix, speech enhancement, CNN, LSTM, time domain †address: 1 Department of Electronic and Telecommunication Engineering, University of Moratuwa, Sri Lanka 2 School of Engineering, The Australian National University, Australia †email: yalegamammoab.20@uom.lk, wageesham@uom.lk 1 Introduction Blind estimation of the relative transfer function (ReTF) without requiring positional knowledge of receivers or sound sources is highly attractive in practical applications such as robot audition [22], drone audition [17], teleconferencing [2], and hearing aids [20], which involve tasks such as sound source localization, speech enhancement, speaker separation, and acoustic echo cancellation [7, 28, 34, 12]. However, these applications typically involve multiple simultaneously active sound sources, where the assumption of W-disjoint orthogonality [38] required for ReTF estimation does not hold. To overcome this limitation, the relative transfer matrix (ReTM) [1] has been introduced as a generalization of the ReTF for acoustic environments with multiple simultaneously active sound sources and multiple microphones, providing an essential component in audio signal processing techniques and applications. In this paper, we propose machine learning-based methods for ReTM estimation from microphone recordings. Many approaches to estimate ReTF for a single active sound source have been developed [7, 28, 34, 24, 29, 18, 21, 33, 19, 25, 6, 5, 26, 35, 36, 4, 31, 11]. Among them, covariance-based methods have gained attention due to their practical feasibility as well as simplicity [24, 29, 18, 21, 33, 19], whereas, machine learning-based approaches offer superior performance at an increased computational cost [5, 26, 35, 36, 4, 31, 11]. In [1], the concept of ReTM was introduced to relate the signal received between two sets of microphone groups with respect to all active sources present in a room. Similar to the ReTF [28], the ReTM is independent of source signals but dependent on the spatial location of the sources and the environment [1]. Recently, various studies have been proposed to utilize ReTM for speech enhancement and speaker separation in multi-source noisy reverberant environments [13, 10, 15, 14, 16], which have shown significant performance improvement. Traditional ReTM estimation approach exploits statistical relationships between two microphone groups using covariance matrices [1]. In contrast to the baseline method that learns a spatial mapping between receiver groups, alternative approaches remain largely unexplored. In this paper, we introduce three fully supervised machine learning frameworks for ReTM estimation. We use both time and short-time Fourier transform (STFT) domains, deep learning approaches. All three models assume a stationary environment that all sound sources are unmoving. The contributions of this work are as follows: 1) We propose: i) the STFT Convolutional Network (SCoNet) that uses depthwise convolution; i) the Convolutional Filter and Summation Network (FuSNet) that applies a set of individual convolutional filters; and i) the Long Short-Term Memory (LSTM)-based Autoencoder Network (LAeNet) that uses a shared bidirectional-LSTM followed by a feedforward network. 2) We evaluate ReTM estimation accuracy with respect to the covariance-based approach using five qualitative metrics, signal-to-distortion ratio (SDR), mean square error (MSE), log spectral distortion (LSD), average phase distortion (APD), and relative spectrum error (RSE). 3) We demonstrate the merits of the proposed models in speech denoising. The simulation results are comparable to the baseline method overall, while the STFT domain models consistently outperform it. 2 Problem formulation Consider a reverberant environment with concurrently active ℒL sound sources. In the short time Fourier transform (STFT) domain, we denote Sℓ(f,t)S_ (f,t), ℓ=1,⋯,ℒ =1,·s,L as the source signals. Let there be Q arbitrary distributed microphones in the room. We divide them to two groups of microphones, A\A\ and B\B\ with QAQ_A and QBQ_B microphones, respectively (Q=QA+QBQ=Q_A+Q_B). We denote A(f,t)M_A(f,t) and B(f,t)M_B(f,t) as the vector of received signals at microphone groups A and B, respectively. Then the received signals at each microphone group in matrix form as A(f,t)=A(f)(f,t),M_A(f,t)=H_A(f)S(f,t), (1) B(f,t)=B(f)(f,t),M_B(f,t)=H_B(f)S(f,t), where (f,t)=[S1(f,t),…,Sℒ(f,t)]TS(f,t)=[S_1(f,t),…,S_L(f,t)]^T, and ⋅T\·\^T is the matrix transpose. Here, A(f)∈ℂQA×ℒH_A(f) ^Q_A×L and B(f)∈ℂQB×ℒH_B(f) ^Q_B×L are the matrices with elements defined by the acoustic transfer functions. Note that we consider thermal microphone noise to be negligible. The ReTM, ℛAB(f) R_AB(f), is defined as in [1] ℛAB(f)≜A(f)B(f)†, R_AB(f) _A(f)H_B(f) , (2) where (⋅)†(·) is Moore-Penrose inverse, assuming the validity, i.e., QB≥ℒQ_B . Thus, we can relate the received signal at group A\A\ and B\B\ using A(f,t)= ℛAB(f)B(f,t).M_A(f,t)= R_AB(f)M_B(f,t). (3) In time domain Equation 3 is given by mAj(t)=∑i=1QBrABji(t)∗mBi(t),j=1,⋯,QA,m_A^j(t)= _i=1^Q_Br_AB^ji(t) m_B^i(t), j=1,·s,Q_A, (4) where mAjm_A^j, and mBim_B^i denote the signals of the jthj^th microphone in group A\A\, and the ithi^th microphone in group B\B\, respectively, rABjir_AB^ji denotes the inverse Fourier transform of the j,ith\j,i\^th element of AB(f) R_AB(f). The aim of this paper is to exploit the spatial properties of sound sources by modeling the ReTM with machine learning algorithms in the time domain and STFT domain as fB(⋅)f_B(·), and FB(⋅)F_B(·), respectively, such that MA(f,t)=FB( MB(f,t)), M_A(f,t)=F_B( M_B(f,t)), (5) A(t)=fB(B(t)). m_A(t)=f_B( m_B(t)). (6) The next section proposes machine learning approaches for the modeling of fB(⋅)f_B(·) and FB(⋅)F_B(·). 3 Proposed ReTM estimation models In this section, we present two convolutional networks and one Long Short-Term Memory (LSTM)-based recurrent neural network for the modeling of the ReTM. 3.1 STFT convolutional network (SCoNet) We model FB(⋅)F_B(·) using SCoNet, which applies depthwise convolution operations in the STFT domain, as shown in Figure 1. We input B(f,t) M_B(f,t), STFT of microphone signals in group B\B\, and train to estimate the microphone signals in group A\A\. SCoNet stacks the real and imaginary parts along the channel dimension and performs two-dimensional depthwise convolution across the time and channel axes, allowing the model to learn distinct parameter sets for each frequency component of the relative transfer matrix. The channel-wise outputs of group A\A\ are then stacked to form the STFT representation of MA(f,t)M_A(f,t). Figure 1: Model architecture of SCoNet, where F and T are the number of frequency bins and time frames, respectively. 3.2 Convolutional filter and summation network (FuSNet) Motivated by the effectiveness of convolutional neural networks (CNNs) in modeling the ReTF [4], we propose FuSNet, which parameterizes fB(⋅)f_B(·) using QA×QBQ_A× Q_B learnable one-dimensional convolutional filters, whose weights correspond to rABji(t)r_AB^ji(t) in Equation 4. Figure 3 shows the FuSNet architecture. We input non-overlapping segments of length L, and a context window of size 3L3L from the time frames of microphone signals at group B\B\. The context window is chosen such that the convolution output matches the original segment length, setting the filter size to L. Note that the window length should be longer than the room’s reverberation time to satisfy the multiplicative transfer function [4, 3]. The convolution outputs are regrouped and summed to obtain the time domain microphone signals at group A\A\. Figure 2: Model architecture of FuSNet for the case of QB=4Q_B=4 and QA=3Q_A=3. Figure 3: Model architecture of LAeNet with a shared BiLSTM. 3.3 LSTM based autoencoder network (LAeNet) We propose LAeNet, an LSTM-based autoencoder for modeling of FB(⋅)F_B(·), as shown in Figure 3. This design exploits LSTMs’ ability to capture narrow-band spatial information [37]. In LAeNet, each frequency bin of the concatenated STFTs A(f,t) M_A(f,t) and B(f,t) M_B(f,t) is independently processed by a bidirectional-LSTM (BiLSTM) layer with shared weights. The resulting temporal sequences are passed through layer normalization to stabilize and accelerate training, then averaged over time to yield feature vectors of size 4(QB+QA)4(Q_B+Q_A), where the increase in dimensionality stems from the BiLSTM. Each feature vector is subsequently mapped by a fully connected network k(⋅)k(·) to estimate the ReTM coefficients, which are finally used to estimate the group A\A\ microphone signals A(f,t) M_A(f,t) as in Equation 3. 3.4 Loss function and training To improve the stability and performance of the proposed models, we train using a weighted sum of the negative signal-to-distortion ratio (SDR) computed in the time domain, and the relative spectrum error (RSE) measured in the STFT domain as: L=−αLSDR(A,^A)+βLRSE(A,^A),L=-α L_SDR( m_A, m_A)+β L_RSE( M_A, M_A), (7) where ^A m_A, and ^A M_A denote the estimated microphone signals at group A\A\ in the time and STFT domains, respectively. This joint optimization enforces consistency across both representations and improves robustness to noise and artifacts that may manifest differently in the time and frequency domains. Here, α and β are weight factors for LSDR=1T∑t=1T10log10(|mA(t)|22|mA(t)−m^A(t)|22)L_SDR= 1T _t=1^T10 _10 ( |m_A(t)|_2^2|m_A(t)- m_A(t)|_2^2 ) (8) and LRSE=1FT∑t=1T∑f=1F10log10(|M^A(f,t)−MA(f,t)|2|MA(f,t)|2),L_RSE= 1FT _t=1^T _f=1^F10 _10 ( | M_A(f,t)-M_A(f,t)|^2|M_A(f,t)|^2 ), (9) respectively. For training, we employ the Adam optimizer [9] with a learning rate that is adaptively reduced over epochs based on the validation set performance. 4 Experiments 4.1 Experimental methodology We utilize an open-source toolbox [8] to model the room impulse response from the sound sources to irregularly distributed microphones in a 6×7×36× 7× 3 m rectangular room (T60=500T_60=500 ms). We assign Q=7Q=7 with QA=3Q_A=3, and QB=4Q_B=4 number of receivers to group A\A\ and B\B\, respectively. We consider four scenarios. A1: two white Gaussian noise (WGN) sources, A2: two noise sources (air conditioner noise, music), and B: three sources (speech, air conditioner noise, music). Scenario C uses a total of Q=12Q=12 microphones with QA=5Q_A=5, and QB=7Q_B=7, drives with the same set of sources as B. The received signals are added with 4040 dB SNR of WGN at each microphone and down-sampled to 1616 kHz. For Scenario A1, the training, validation, and test recordings are 3 minutes, 1 minute, and 1 minute, respectively. For all other scenarios, 50 s and 10 s recordings are used for training and testing, respectively. For the model hyperparameter configuration, FuSNet is designed with a window and context size of 8192 samples. For SCoNet and LAeNet, the input recordings are converted to STFT frames using a Hann window of size 8192 with 50% overlap. The window length is selected based on the room’s reverberation time to ensure that it is sufficiently long for the multiplicative transfer function assumption to hold [3]. The weight factors α and β are set to 1 and 10, respectively. For a comprehensive evaluation of the proposed models, we adopt the covariance-based approach [1] as the baseline. To fairly assess performance across all scenarios, we use five quantitative metrics: (i) SDR, (i) mean square error (MSE), defined as MSE=10log10(1T∑t=1T|mA(t)−m^A(t)|2),MSE=10 _10 ( 1T _t=1^T|m_A(t)- m_A(t)|^2 ), (i) log spectral distortion (LSD), defined as LSD=1N∑η=1N(10log10|MA(fη,t)|2/|M^A(fη,t)|2),2LSD= 1N _η=1^N (10 _10|M_A(f_η,t)|^2/| M_A(f_η,t)|^2 ),^2 (iv) average phase distortion (APD), defined as APD=1N∑η=1N|∠MA(fη,t)−∠M^A(fη,t)|,APD= 1N _η=1^N | M_A(f_η,t)- M_A(f_η,t) |, and (v) RSE. 4.2 Results and discussion Table 1: ReTM estimation accuracy for various scenarios (Channel 1/Average). Scenario Method SDR (dB)↑ MSE (dB)↓ LSD (dB)↓ APD (rad)↓ RSE (dB)↓ A1 Baseline 26.56/26.126.56/26.1 −43.84/−43.56-43.84/-43.56 1.21/1.191.21/1.19 0.1/0.10.1/0.1 −24.76/−24.10-24.76/-24.10 SCoNet 24.76/24.2424.76/24.24 −40.88/−40.54-40.88/-40.54 0.89/0.960.89/0.96 0.13/0.140.13/0.14 −25.62/−24.89-25.62/-24.89 FuSNet 29.13/28.46 29.13/28.46 −59.25/−58.75 -59.25/-58.75 0.5/0.49 0.5/0.49 0.08/0.09 0.08/0.09 −48.61/−48.4 -48.61/-48.4 LAeNet 25.12/24.0925.12/24.09 −55.24/−54.38-55.24/-54.38 1.03/1.131.03/1.13 0.12/0.130.12/0.13 −20.25/−19.74-20.25/-19.74 A2 Baseline 21.2/20.521.2/20.5 −36.25/−36.57-36.25/-36.57 2.02/2.032.02/2.03 0.21/0.210.21/0.21 −25.28/24.65-25.28/24.65 SCoNet 26.26/26.5926.26/26.59 −43.08/−42.47-43.08/-42.47 1.22/1.23 1.22/1.23 0.18/0.19 0.18/0.19 −24.61/−24.06-24.61/-24.06 FuSNet 30.96/29.62 30.96/29.62 −55.2/−54.94 -55.2/-54.94 2.47/2.592.47/2.59 0.23/0.20.23/0.2 −39.81/−33.78 -39.81/-33.78 LAeNet 28.07/26.2128.07/26.21 −52.32/−51.53-52.32/-51.53 1.68/1.701.68/1.70 0.22/0.210.22/0.21 −18.17/−17.67-18.17/-17.67 B Baseline 16.1/16.2916.1/16.29 −36.65/−37.49-36.65/-37.49 2.48/2.52.48/2.5 0.23/0.240.23/0.24 −17.39/−16.85-17.39/-16.85 SCoNet 22.42/21.96 22.42/21.96 −42.78/−42.98-42.78/-42.98 1.53/1.54 1.53/1.54 0.19/0.19 0.19/0.19 −20.76/−20.33 -20.76/-20.33 FuSNet 22.51/21.9122.51/21.91 −45.12/−45.17 -45.12/-45.17 4.89/4.964.89/4.96 0.71/0.720.71/0.72 −19.60/−19.16-19.60/-19.16 LAeNet 22.36/21.8822.36/21.88 −45.01-45.01 / −45.18 -45.18 1.78/1.881.78/1.88 0.2/0.210.2/0.21 −17.58/−17.21-17.58/-17.21 C Baseline 23.05/23.2523.05/23.25 −43.6/−43.35-43.6/-43.35 1.03/1.09 1.03/1.09 0.15/0.150.15/0.15 −34.03/−33.22-34.03/-33.22 SCoNet 29.59/27.8529.59/27.85 −44.12/−43.27-44.12/-43.27 1.06/1.171.06/1.17 0.14 0.14 / 0.160.16 −28.63/−27.43-28.63/-27.43 FuSNet 40.33/38.88 40.33/38.88 −62.89/−61.98 -62.89/-61.98 1.95/2.101.95/2.10 0.15/0.160.15/0.16 −34.07/−33.27 -34.07/-33.27 LAeNet 29.35/28.5729.35/28.57 −52/−51.77-52/-51.77 1.39/1.491.39/1.49 0.14/0.15 0.14/0.15 −20.15/−19.97-20.15/-19.97 The ReTM estimation accuracy results are given in Table 1. From the results in scenario A1, we observe that FuSNet achieves the highest performance across both time and frequency domain metrics. This improvement of FuSNet can be attributed to the use of WGN sources that provide a uniform spectral distribution and thereby facilitate more accurate frequency domain estimation. The other proposed methods, SCoNet and LAeNet, also achieve high estimation accuracy, performing comparably to the baseline method. In scenario A2, FuSNet again achieves the highest overall performance across most evaluation metrics, with noticeable degradation in LSD and APD. The key difference between scenarios A1 and A2 lies in the replacement of WGN with specific noise sources, which exhibit non-uniform frequency distributions. This indicates that FuSNet performs best with uniform spectra but degrades under non-uniform conditions. Moreover, SCoNet yields performance on par with LAeNet and surpasses the state-of-the-art method, excluding the RSE score, where the baseline approach shows the second-best performance. From scenarios B and C, we observe that all four methods exhibit improved ReTM estimation accuracy as the number of microphones in both groups increases, consistently across all five evaluation metrics. The traditional method exhibits significantly reduced accuracy, and is seen to be the worst performing method among all methods, except for the LSD and RSE measures. Compared to the baseline, both SCoNet and FuSNet, show substantially superior performance. SCoNet performs best in scenario B with an MSE score lower than FuSNet. Whereas, FuSNet achieves the highest performance in scenario C, excluding the LSD and APD measures. Overall, it can be verified that although the baseline method achieves satisfactory results, it is clearly outperformed by the proposed deep learning-based models, particularly in the time domain measures, e.g., MSE and SDR. Table 2: Computational complexity of each proposed model. Method QA=3,QB=4QA=3,Q_B=4 QA=5,QB=7Q_A=5,Q_B=7 Param. Latency Param. Latency SCoNet 196.7196.7k 6.77ms 573.6573.6k 6.61ms FuSNet 98.398.3k 1.07ms 286.8286.8k 2.41ms LAeNet 263.5263.5k 1.80s 652.5652.5k 1.78s Table 2 compares the computational efficiency of the proposed models in terms of the parameter count, alongside the inference latency measured on an NVIDIA GeForce RTX 3090 GPU. Across both scenarios, FuSNet demonstrates the lowest latency and the smallest parameter count, whereas LAeNet exhibits the highest latency and the largest number of parameters. Notably, FuSNet and SCoNet benefit from faster training due to their comparatively low latency, although their memory usage increases linearly with the window length. In contrast, LAeNet maintains nearly constant memory usage as the window length increases, owing to its shared LSTM and feedforward network architecture. This design allows efficient parameter utilization for longer windows but results in substantially longer training times due to higher latency. 5 Application into speech enhancement Motivated by the computational efficiency of the proposed architectures, in this section we further evaluate the ReTM estimation accuracy in a speech denoising application [10]. Denote ℛAB(N) R_AB^(N) be the noise sources ReTM. In brief, [10] proposed to enhance the speech from the noisy speech recordings with the known ℛAB(N) R_AB^(N), which can be estimated using the noise-only recordings, as ^≜A− ℛAB(N)B=[A− ℛAB(N)B]S, S _A- R_AB^(N)M_B= [h_A- R_AB^(N)h_B ]S, (10) where Ah_A and Bh_B be the acoustic transfer function vectors from the speech source to group A\A\ and B\B\ microphones, respectively. We note that S is a QA×1Q_A× 1 vector consists QAQ_A copies of estimated target speech signal S. We evaluate speech denoising performance in scenarios B and C (Sec. 4), each containing a single speech source, where the ReTM is estimated from 1-minute noise-only recordings. For analysis we use the SDR [30], and the Short-Time Objective Intelligibility (STOI) [27]. Table 3 depicts the speech enhancement accuracy for both B and C scenarios. The results confirm that enhanced performance (+0.5+0.5 dB in SDR, 13%13\% in STOI) for all methods. Although FuSNet significantly outperforms ReTM estimation on most metrics, its speech denoising performance degrades, confirming the advantage of time-frequency domain ReTM estimation over the time domain approach. Compared to FuSNet, both LAeNet and SCoNet, demonstrate significantly superior performance. LAeNet achieves the highest performance in both scenarios B and C, with an average SDR of +10.55+10.55 dB and an average STOI of 37%37\%. Whereas, SCoNet ranks second in both scenarios, yielding an average SDR of +8.41+8.41 dB and an average STOI 33%33\%. In contrast, the traditional method demonstrates comparable performance, although the proposed models, except FuSNet, consistently achieve better results in both scenarios. From informal listening, we find that FuSNet’s denoised speech remains significant echo noise, leading to lower performance, which could be improved through dereverberation algorithms [23, 32] that we will investigate in future work. The audio samples of this work can be found on GitHub.11 1 https://github.com/oshanyalegama/Denoised_ReTM_DL Table 3: Evaluation results on the scenarios B & C. Method B C SDR STOI SDR STOI Noisy −2.70-2.70 0.540.54 −2.70-2.70 0.540.54 Baseline 6.066.06 0.870.87 4.074.07 0.850.85 SCoNet 6.456.45 0.870.87 4.964.96 0.870.87 FuSNet 2.212.21 0.800.80 −1.19-1.19 0.670.67 LAeNet 8.67 8.67 0.92 0.92 7.03 7.03 0.91 0.91 6 Conclusion We proposed SCoNet, FuSNet, and LAeNet to estimate the ReTM using machine learning. Our proposed approaches use either time or time-frequency domains to model the inter-group spatial relationship. Experimental results confirmed that FuSNet consistently outperformed other approaches in terms of accuracy, while SCoNet and LAeNet achieved performance comparable to the baseline method. We further validated the effectiveness of the proposed methods in a speech denoising application. In future work, we aim to extend these models to source separation, speech dereverberation, and investigate optimal microphone grouping strategies for ReTM-based applications. 7 Acknowledgements This work was supported by the Accelerating Higher Education Expansion and Development (AHEAD) operation (Grant No. 6026-LK/8743-LK, World Bank) and a data scholarship from the Linguistic Data Consortium, University of Pennsylvania. The authors thank Prof. Thushara Abhayapala for fruitful discussions and support on this work. 8 Use of Generative AI Disclosure Generative AI tools were used to assist in drafting and refining portions of the scripts. Additionally, grammar-checking models were utilized to improve language clarity, correctness, and overall readability. References [1] T. D. Abhayapala, L. Birnie, M. Kumar, D. Grixti-Cheng, and P. N. Samarasinghe (2023) Generalizing the relative transfer function to a matrix for multiple sources and multichannel microphones. In Proc.Eur. Signal Process. Conf., p. 336–340. Cited by: §1, §1, §2, §4.1. [2] S. Araki, M. Fujimoto, K. Ishizuka, H. Sawada, and S. Makino (2008) Speaker indexing and speech enhancement in real meetings/conversations. In Proc.IEEE Int. Conf. on Acoust., Speech and Signal Process., p. 93–96. Cited by: §1. [3] Y. Avargel and I. Cohen (2007) On multiplicative transfer function approximation in the short-time fourier transform domain. IEEE Signal Process. Letters 14 (5), p. 337–340. Cited by: §3.2, §4.1. [4] L. Birnie, P. Samarasinghe, T. Abhayapala, and D. Grixti-Cheng (2021) Noise retf estimation and removal for low snr speech enhancement. In ’IEEE Workshop Mach. Learning Signal Process.’, p. 1–6. Cited by: §1, §3.2. [5] A. Brendel, J. Zeitler, and W. Kellermann (2022) Manifold learning-supported estimation of relative transfer functions for spatial filtering. In Proc.IEEE Int. Conf. on Acoust., Speech and Signal Process., p. 8792–8796. Cited by: §1. [6] I. Cohen (2004) Relative transfer function identification using speech signals. IEEE Trans. on Speech and Audio Process. 12 (5), p. 451–459. Cited by: §1. [7] S. Gannot, D. Burshtein, and E. Weinstein (2001) Signal enhancement using beamforming and nonstationarity with applications to speech. IEEE Trans. on Signal Process. 49 (8), p. 1614–1626. Cited by: §1, §1. [8] E. A. Habets (2006) Room impulse response (RIR) generator. Note: [Online]. Available: https://w.audiolabserlangen.de/fau/professor/habets/software/rir-generator Cited by: §4.1. [9] D. P. Kingma and J. Ba (2014) Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980. Cited by: §3.4. [10] M. Kumar, L. Birnie, T. Abhayapala, S. A. Holzinger, A. Bastine, D. Grixti-Cheng, and P. Samarasinghe (2024) Speech denoising in multi-noise source environments using multiple microphone devices via relative transfer matrix. In Proc.Eur. Signal Process. Conf., p. 336–340. Cited by: §1, §5, §5. [11] D. Levi, A. Sofer, and S. Gannot (2025) PeerRTF: robust mvdr beamforming using graph convolutional network. IEEE Trans. on Audio, Speech, and Lang. Process.. Cited by: §1. [12] W. N. Manamperi, T. D. Abhayapala, L. Brinie, J. Zhang, and P. N. Samarasinghe (2025) Drone audition: on measurements and modeling of drone-related transfer functions. IEEE/ACM Trans. on Audio, Speech, and Lang. Process. 33, p. 1775 – 1786. Cited by: §1. [13] W. N. Manamperi and T. D. Abhayapala (2024) Relative transfer matrix for drone audition applications: source enhancement. In Proc.Asia-Pacific Signal and Inf. Process. Assoc. Annu. Summit and Conf., p. 1–6. Cited by: §1. [14] W. N. Manamperi and T. D. Abhayapala (2024) Successive speaker relative transfer function estimation through relative transfer matrix in noisy reverberant environments. In Proc.Asia-Pacific Signal and Inf. Process. Assoc. Annu. Summit and Conf., p. 1–6. Cited by: §1. [15] W. N. Manamperi and T. D. Abhayapala (2025) Relative transfer matrix estimator using covariance subtraction. arXiv preprint arXiv:2510.19439. Cited by: §1. [16] W. N. Manamperi (2026) Multiple speaker separation in reverberant rooms under low snr conditions using the relative transfer matrix. accepted for publication in Appl. Acoust.. Cited by: §1. [17] W. Manamperi, T. D. Abhayapala, J. (. Zhang, and P. N. Samarasinghe (2022) Drone audition: Sound source localization using on-board microphones. IEEE/ACM Trans. on Audio, Speech, and Lang. Process. 30, p. 508 – 519. Cited by: §1. [18] S. Markovich, S. Gannot, and I. Cohen (2009) Multichannel eigenspace beamforming in a reverberant noisy environment with multiple interfering speech signals. IEEE/ACM Trans. on Audio, Speech, and Lang. Process. 17 (6), p. 1071–1086. Cited by: §1. [19] S. Markovich-Golan, S. Gannot, and W. Kellermann (2018) Performance analysis of the covariance-whitening and the covariance-subtraction methods for estimating the relative transfer function. In Proc.Eur. Signal Process. Conf., p. 2499–2503. Cited by: §1. [20] D. Marquardt, E. Hadad, S. Gannot, and S. Doclo (2016) Incorporating relative transfer function preservation into the binaural multi-channel wiener filter for hearing aids. In Proc.IEEE Int. Conf. on Acoust., Speech and Signal Process., p. 6500–6504. Cited by: §1. [21] W. Middelberg, H. Gode, and S. Doclo (2023) Relative transfer function vector estimation for acoustic sensor networks exploiting covariance matrix structure. In Proc.IEEE Workshop on Applications of Signal Process. to Audio and Acoust., p. 1–5. Cited by: §1. [22] K. Nakadai, T. Lourens, H. G. Okuno, and H. Kitano (2000) Active audition for humanoid. In Natl. Conf. Artif. Intell., p. 832–839. Cited by: §1. [23] P. A. Naylor and N. D. Gaubitch (2010) Speech dereverberation. Springer. Cited by: §5. [24] R. Serizel, M. Moonen, B. Van Dijk, and J. Wouters (2014) Low-rank approximation based multichannel wiener filter algorithms for noise reduction with application in cochlear implants. IEEE/ACM Trans. on Audio, Speech, and Lang. Process. 22 (4), p. 785–799. Cited by: §1. [25] O. Shalvi and E. Weinstein (1996) System identification using nonstationary signals. IEEE Trans. on Signal Process. 44 (8), p. 2055–2063. Cited by: §1. [26] A. Sofer, T. Kounovskỳ, J. Čmejla, Z. Koldovskỳ, and S. Gannot (2021) Robust relative transfer function identification on manifolds for speech enhancement. In Proc.Eur. Signal Process. Conf., p. 401–405. Cited by: §1. [27] C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen (2011) An algorithm for intelligibility prediction of time-frequency weighted noisy speech. IEEE/ACM Trans. on Audio, Speech, and Lang. Process. 19 (7), p. 2125–2136. Cited by: §5. [28] R. Talmon, I. Cohen, and S. Gannot (2009) Relative transfer function identification using convolutive transfer function approximation. IEEE Trans. on Audio, Speech, and Lang. Process. 17 (4), p. 546–555. Cited by: §1, §1, §1. [29] R. Varzandeh, M. Taseska, and E. A. P. Habets (2017) An iterative multichannel subspace-based covariance subtraction method for relative transfer function estimation. In Proc.Joint Workshop Hands-free Speech Comm. and Microphone Arrays, p. 11–15. Cited by: §1. [30] E. Vincent, R. Gribonval, and C. Févotte (2006) Performance measurement in blind audio source separation. IEEE Trans. on Audio, Speech, and Lang. Process. 14 (4), p. 1462–1469. Cited by: §5. [31] Z. Wang and D. Wang (2018) Mask weighted stft ratios for relative transfer function estimation and its application to robust asr. In Proc.IEEE Int. Conf. on Acoust., Speech and Signal Process., p. 5619–5623. Cited by: §1. [32] Z. Wang and D. Wang (2020) Deep learning based target cancellation for speech dereverberation. IEEE/ACM Trans. on Audio, Speech, and Lang. Process. 28, p. 941–950. Cited by: §5. [33] E. Warsitz and R. Haeb-Umbach (2007) Blind acoustic beamforming based on generalized eigenvalue decomposition. IEEE Trans. on Audio, Speech, and Lang. Process. 15 (5), p. 1529–1539. Cited by: §1. [34] E. Warsitz, A. Krueger, and R. Haeb-Umbach (2008) Speech enhancement with a new generalized eigenvector blocking matrix for application in a generalized sidelobe canceller. In Proc.IEEE Int. Conf. on Acoust., Speech and Signal Process., p. 73–76. Cited by: §1, §1. [35] B. Yang, X. Li, and H. Liu (2021) Supervised direct-path relative transfer function learning for binaural sound source localization. In Proc.IEEE Int. Conf. on Acoust., Speech and Signal Process., p. 825–829. Cited by: §1. [36] B. Yang, H. Liu, and X. Li (2021) Learning deep direct-path relative transfer function for binaural sound source localization. IEEE/ACM Trans. on Audio, Speech, and Lang. Process. 29, p. 3491–3503. Cited by: §1. [37] Y. Yang, C. Quan, and X. Li (2023) McNet: fuse multiple cues for multichannel speech enhancement. In Proc.IEEE Int. Conf. on Acoust., Speech and Signal Process., p. 1–5. Cited by: §3.3. [38] O. Yilmaz and S. Rickard (2004) Blind separation of speech mixtures via time-frequency masking. IEEE Signal Process. Mag. 52 (7), p. 1830–1847. Cited by: §1.