Paper deep dive
SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios
Ziyang Jiang, Yu Chen, Zexu Pan, Xinyuan Qian, Bowen Xing, Ivor W. Tsang, Xu-Cheng Yin, Haizhou Li
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 7/5/2026, 12:55:53 PM
Summary
SelectTSL is an end-to-end architecture designed for prompt-guided selective target sound localization in complex, multi-source acoustic environments. Unlike conventional Sound Source Localization (SSL) which is semantic-blind, or Target Sound Extraction (TSE) which often degrades spatial information, SelectTSL uses multimodal prompts (text or audio) to selectively localize a specific target. The system employs a Prompt-Guided Selective Attention (PGSA) module to generate extraction-informed embeddings (EIEs) that guide an Inter-channel Phase Difference (IPD) enhancer. This allows the model to jointly estimate the Direction of Arrival (DoA) and the target-source cardinality (the number of active target sources), even when the number of targets varies over time.
Entities (6)
Relation Signals (4)
SelectTSL → contains → Prompt-Guided Selective Attention (PGSA)
confidence 100% · Specifically, we design a target-aware selective localization strategy that employs a Prompt-Guided Selective Attention (PGSA) module
SelectTSL → contains → Progressive Refinement Temporal Modeling Module (PRTM)
confidence 100% · Progressive Refinement Temporal Modeling Module (PRTM) DoA Estimator
PGSA → generates → Extraction-Informed Embeddings (EIEs)
confidence 100% · The PGSA module outputs the target magnitude spectrogram and a stack of EIEs
CLAP → usedfor → Prompt Encoding
confidence 100% · We employ a frozen CLAP model to extract audio cue x cue (t) and text x text
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Humans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remains challenging for current deep learning-based systems. Sound source localization (SSL) has achieved remarkable success with deep learning, yet most methods localize all active sources without selectivity. Conversely, target sound extraction (TSE) extracts sources using multimodal prompts but typically fails to preserve the multichannel spatial information required for accurate localization. To bridge this gap, we formulate the task of prompt-guided selective target sound localization and propose SelectTSL, an end-to-end architecture that localizes only the user-specified target in multi-source acoustic scenes. Specifically, we design a target-aware selective localization strategy that employs a Prompt-Guided Selective Attention Module (PGSA) to generate prompt-informed embeddings. These embeddings guide an inter-channel phase difference (IPD) enhancer to refine raw phase cues, fusing with target magnitudes to jointly estimate direction of arrival (DoA) and target-source cardinality, i.e., the number of target sound sources. This coupled design effectively focuses on the user-specified target spatial cues for selective localization and also handles time-varying numbers of target sources. Extensive experiments on both synthetic data and real-world recordings demonstrate that our proposed method consistently outperforms other baselines and exhibits robust generalization to real acoustic environments.
Tags
Links
- Source: https://arxiv.org/abs/2607.02343v1
- Canonical: https://arxiv.org/abs/2607.02343v1
Trouble viewing inline? Open PDF directly →
Full Text
85,331 characters extracted from source content.
Expand or collapse full text
1 SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios Ziyang Jiang, Student Member, IEEE, Yu Chen, Zexu Pan, Member, IEEE, Xinyuan Qian, Senior Member, IEEE, Bowen Xing, Ivor W. Tsang, Fellow, IEEE, Xu-Cheng Yin, Senior Member, IEEE, Haizhou Li, Fellow, IEEE Abstract—Humans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remains challenging for current deep learn- ing based systems. Sound source localization (SSL) has achieved remarkable success with deep learning, yet most methods localize all active sources without selectivity. Conversely, target sound extraction (TSE) extracts sources using multimodal prompts but typically fails to preserve the multichannel spatial information re- quired for accurate localization. To bridge this gap, we formulate the task of prompt-guided selective target sound localization and propose SelectTSL, an end-to-end architecture that localizes only the user-specified target in multi-source acoustic scenes. Specifi- cally, we design an target-aware selective localization strategy that employs a Prompt-Guided Selective Attention Module (PGSA) to generate prompt-informed embeddings. These embeddings guide an inter-channel phase difference (IPD) enhancer to refine raw phase cues, fusing with target magnitudes to jointly estimate direction of arrival (DoA) and target-source cardinality (i.e., the number of target sound sources). This coupled design effectively focuses on the user-specified target spatial cues for selective local- ization and also handles time-varying numbers of target sources. Extensive experiments on both synthetic data and real-world recordings demonstrate that our proposed method consistently outperforms other baselines and exhibits robust generalization to real acoustic environments. Dataset and code will be released. I. INTRODUCTION S OUND source localization (SSL), which estimates the spatial location of acoustic events, supports a wide range of array-based audio and speech applications. For instance, in smart speakers, SSL enables direction-aware target speech enhancement and far-field automatic speech recognition (ASR) by steering microphone-array beamformers [1]. In hearing aids, it supports directional noise reduction via spatial filtering [2]. Despite remarkable progress of SSL via signal process- ing (e.g., generalized cross-correlation with phase transform (GCC-PHAT) [3], multiple signal classification (MUSIC) [4]) or deep neural networks (e.g., convolutional recurrent neu- ral networks (CRNNs) [5]), existing methods are inherently semantic-blind, treating all acoustic sources as generic signals without identity awareness: all active sources are localized without selectivity, including interfering sounds. This differs from human auditory perception where listeners can selec- tively attend to a target sound when competing speakers and background noise exist, i.e., the “cocktail party problem” [6]. As illustrated in Fig. 1(a), conventional SSL systems indiscriminately localize both target and irrelevant sources (e.g., dog barks and noise), yielding non-selective direction-of- arrival (DoA) trajectories and preventing users from directing the system to focus on a specific target such as speech. Target (Speech) Interference (Dog barks) Noise Target Trajectory Interference Trajectory Dual-channel microphone Conventional SSL (Semantic-Blind) Localizes ALL sources Output: Non-selective Target (Speech) Interference (Dog barks) Noise Target Trajectory Interference Trajectory Ours (SelectTSL: Prompt-guided selective) Localizes ONLY the target source Output: Selective SSL Prompt Input Another Speech TSE - DOAnet SelectiveDoA Estimation “locate the speech” (a)(b) Fig. 1. Illustration of conventional semantic-blind SSL and our proposed prompt-guided SelectTSL: (a) Conventional SSL localizes all active sources, yielding non-selective DoA trajectories, whereas (b) SelectTSL uses a text and audio prompt (“locate the speech”) to focus on the target speech which only provides its corresponding DoA trajectory. To associate semantics with location, sound event localiza- tion and detection (SELD) methods [7], [7], [8] detect and localize all events of known classes. However, they perform passive scene analysis that lacks an interactive mechanism to filter sources based on user intent [9]. Conversely, in the field of target sound extraction (TSE), methods [10]–[13] extract specific sound sources guided by audio or text prompts. Nevertheless, TSE focuses more on waveform reconstruction process rather than spatial sensing, thus spatial cues like inter- channel phase difference (IPD) are often degraded for precise SSL. This leads to a fundamental mismatch: TSE is aware of what to extract but not where to localize, whereas conventional SSL knows where sources are but not which semantic target to follow. Therefore, it remains an open problem to build a unified framework that is both semantic-aware and location- aware. In this paper, we propose the SelectTSL network. Unlike standard approaches, SelectTSL is an end-to-end localiza- tion framework, where prompt-guided target extraction guides subsequent spatial estimation. In particular, a Prompt-Guided Selective Attention (PGSA) module, conditioned by multi- modal prompts (text or audio), acts as a semantic filter to purify target-consistent features from the audio mixture. These features are then fused with spatial cues (e.g., IPD) and fed into a dedicated DoA estimator. This design decouples semantic selection from spatial estimation, enabling robust localization of user-specified targets even in noisy and multi- source environments. Our contributions are listed as: 1) We formulate a prompt-guided selective target sound arXiv:2607.02343v1 [cs.SD] 2 Jul 2026 2 localization task in dynamic, noisy and multi-source environments, where sources are moving and the number of active target sources is unknown and varies over time. The goal is to localize only the user-specified target(s) while suppressing interference. 2) We propose SelectTSL, an end-to-end framework that supports both text and audio prompts. A prompt-guided PGSA module produces extraction-informed embed- dings (EIEs) that condition an IPD Enhancer to refine spatial phase cues, which are fused with target magni- tudes for DoA estimation. 3) We handle unknown and time-varying active targets via a lightweight cardinality head that jointly predicts DoA heatmaps and frame-level target-source counts, enabling stable localization and tracking under moving speakers and intermittent activity. 4) Extensive experiments on synthetic and real data demon- strate consistent improvements over strong baselines in localization and tracking, and confirm robust generaliza- tion under realistic room acoustics. I. RELATED WORK Conventional SSL approaches excel at estimating spatial direction but are typically prompt-agnostic, localizing all active sources indiscriminately. Conversely, TSE provides controllable target selection via auxiliary cues to reconstruct clean sound. However, it prioritizes signal fidelity over spatial consistency, often distorting the spatial information needed for precise localization. Prompt-guided selective target sound localization bridges the gap between semantic-aware extrac- tion and spatial-aware localization, effectively unifying these two objectives. Despite its promise, existing studies in this emerging field are still limited by unimodal prompting (text or audio only) and restricted evaluation scenarios (e.g., handling at most a single target per query, with limited evaluation in moving-source and noisy environments). The following section reviews literature in TSE, SSL, and prompt-guided localization. The significant methods summarized in Table I. A. Target Sound Extraction (TSE) TSE aims to extract a specific type of sound source (e.g., speech, music, or other acoustic events) from a mixture, conditioned on explicit auxiliary cues. Existing methods have explored auxiliary cues ranging from visual signals [14]–[17], spatial information [18], [19], reference audio [10], [20]–[22] and semantic text descriptions [11], [23]–[25]. While visual cues are highly effective in speech-oriented settings by lever- aging articulatory motions (e.g., lip/face movements) [14], [15], they are not generally applicable to arbitrary acoustic events. Spatial cues used as auxiliary prompts often require additional priors (e.g., array geometry or target direction) beyond the mixture audio, and are thus beyond our scope. Therefore, we focus on text and audio cues in the remainder of this section. For text-prompted extraction, the core mechanism involves mapping a natural language caption (e.g., “piano”, “glass breaking”) into a semantic embedding space, utilizing pre- trained contrastive models such as Contrastive Language- Audio Pretraining (CLAP) [24]. This embedding then acts as a condition to steer the separation network, predicting a time- frequency mask or waveform residual that extracts the source matching the description [11], [25], [26]. In contrast, audio-prompted extraction utilizes a reference audio clip as the enrollment cue. This mechanism generally operates in two distinct modes based on the granularity of the guidance: (i) query by example, where the reference shares the same semantic class as the target but differs in instance (e.g., using a generic dog bark to extract a specific dog) [27]; and (i) target enrollment, commonly used in TSE, which utilizes a clean reference utterance to guide the extraction of the target speaker [10], [20], [22]. Architecturally, these prompt encoders are integrated into separation backbones, such as Conv-TasNet [28], Dual- Path Recurrent Neural Network (DPRNN) [29] and TF- GridNet [30], via varying fusion mechanisms, ranging from simple feature-wise linear modulation (FiLM) to more com- plex cross-attention layers. However, while prompt-guided TSE offers fine-grained controllability over which source to recover, it is primarily optimized for signal reconstruction fidelity (e.g., scale-invariant signal-to-noise ratio (SI-SNR)) rather than spatial consistency. Unless explicitly designed to preserve inter-channel phase consistency, the reconstruction objective of TSE tends to distort the spatial cues essential for localization. In essence, TSE determines what to extract, yet remains agnostic to where the source is located. B. Sound Source Localization (SSL) In contrast to TSE that controls which source to extract, SSL addresses where sources are by estimating DoA from multichannel audio inputs [31]. Recent advances have largely adopted deep learning to map time-frequency (TF) and spatial features to DoA estimates. In parallel, the vision community has studied audio-visual sound source localization in videos, where visual context is used to select and localize the sounding object [32]–[34]. Contemporary systems generally follow a common pipeline: a feature extraction backbone followed by a task-specific head. Despite this shared framework, these methods differ primarily in two aspects: (i) the choice of input representations, ranging from raw short-time Fourier transform (STFT) spectrograms to spatially explicit cues such as IPD, inter-channel level difference (ILD), covariance eigenvectors, and generalized cross correlation (GCC) [7], [35]; and (i) the modeling strategy for these representations, spanning from bandwise processing that preserves narrow-band spatial cues to full-band mechanisms that capture cross-frequency correlations [36], [37]. We organize this subsection along these two aspects to explain the sources of recent performance gains. Typical SSL models map TF features to DoA via classifica- tion or regression heads using deep neural backbones, such as convolutional neural networks (CNNs), CRNNs, Conformer, and Transformer [7], [38]–[41]. Beyond raw STFT features, many systems explicitly encode spatial cues. For example, 3 TABLE I BASELINE TAXONOMY WITH TRACK-SLOT (TRACK-WISE) PRIORITY. CATEGORIES MARKED † ARE NON–TRACK-WISE (NO FIXED-TRACK OUTPUTS). CategoryModelTask familyInput / FeatureTarget Track-wise Multi-ACCDoASELDMch. STFTMulti-ACCDoA IPDNetSSLMch. STFTDoA (via DP-IPD) MIMO-DoAnetSSLMch. STFTMulti-out DoA EINV2SELDFOATrack-wise SELD embed-ACCDoASELDCLAP-aug. feat.ACCDoA DiffTrack-DoASSLMch. STFTDoA + diff. track loss SALSA / SALSA-LiteSELDSALSA / SALSA-LiteDoA + act. SE-ResNetSELDResNet/SE enc.DoA + act. NGCC-SELDSELDNGCC + SELD pipe.DoA + act. CST-FormerSELDCST attentionDoA + act. DCASE25SELDStereo (2-ch) audioDoA + act. SELD+DoA † SWGFormerSELDTransformer mod.DoA + act. SELDnetSELDMel/FOAACCDoA SELDTSELD(T)Transformer SELDDoA + act. Pure DoA † FN-SSLSSLFB+NB fusionDoA (via DP-IPD) (1src) SRP-DNNSSLDP cues + SRP feat.SRP spectrum SALADnetSSLFOAAttn. DoA Prompt-based † KeywordLocSSL (query)Mch. feat. + kw cuekw-cond. spk DoA Class-cond. SELDSELD (query)Mch. feat. + 1hot cls (FiLM) cls-cond. (DoA+act.) Text-Queried SELSEL (query)GCC-PHAT + txt emb.txt-cond. DoA/traj LocSelectSSL (query)2-ch spec. + enroll sp.enroll-cond. spk DoA GCC-SpeakerSSL (query)GCC-PHAT + spk-wtspk DoA Abbrev. Mch.=multichannel; FOA=First-order Ambisonics; FB/NB=full-/narrow-band; STFT=short-time Fourier transform; CLAP=Contrastive Language-Audio Pretraining; ACCDoA=activity-coupled Cartesian DoA; DP-IPD=direct-path inter-channel phase difference; SRP=steered response power; NGCC=neural generalized cross-correlation; SE=squeeze-and-excitation; enc.=encoder; feat.=feature; mod.=module; pipe.=pipeline; act.=activity; diff.=differentiable; aug.=augmented; spk=speaker; spec.=spectrogram; kw=keyword; cls=class; txt=text; traj=trajectory; 1hot=one-hot; wt=weighting; † Non–track-wise (no fixed-track outputs). SALSA augments log-spectrograms with the normalized prin- cipal eigenvector of the spatial covariance [35], and SALSA- Lite uses normalized IPD [42]. Spatial correlation features such as GCC-PHAT and its neural variant neural GCC- PHAT (NGCC-PHAT) are also commonly used in recent SSL systems [3], [43], [44]. In addition to refined inputs, recent work improves robustness by modeling direct-path cues and cross-frequency structure. For example, FN-SSL and IPDNet use full- and narrow-band fusion to estimate direct-path inter- channel phase difference (DP-IPD) and infer DoA (e.g., via template matching) [36], [37]. SRP-DNN learns direct-path delays and phase, and injects them into steered response power (SRP)-style spectra to refine spatial peaks [45]. Moreover, MIMO-DoAnet predicts source-wise pseudo-spectra to reduce reliance on post-hoc heuristics [46]. Beyond DoA estimation, SELD extends the SSL objective by jointly estimating the sound event type. In this case, SELD models can be used as SSL baselines by solely using its localization branch. For instance, ACCDoA and Multi- ACCDoA [47], [48] use activity-coupled Cartesian DoA (AC- CDoA) vectors as regression targets, with auxiliary duplicating permutation invariant training (ADPIT) to handle sound over- laps of the same class. Another line of work adopts explicit track slots with permutation-invariant training. Representa- tive methods such as EIN [49] employ separate attention- based tracks and have been widely adopted in challenge sys- tems [40], [50]. Variants mainly differ in encoders/backbones (e.g., SALADnet [38], SwG-former [39], CST-Former [41]) and feature fusion strategies [51]. The Detection and Classi- fication of Acoustic Scenes and Events (DCASE) 2025 Task 3 baseline further shows an audio-only stereo setting with a CRNN and a Multi-ACCDoA-like output [52]. While the aforementioned methods focus on frame-level DoA estimates, moving-source scenarios require associating predictions over time. For example, SELDT [53] forms tra- jectories by linking frame-wise source activity and DoA es- timates. Differentiable tracking [54] further integrates iden- tity assignment (e.g., Hungarian matching) into training to optimize sequence-level objectives. Although effective for trajectory estimation, such track-based designs inherently rely on a fixed number of output slots and lack a mechanism for user-specified target selection. This limitation motivates prompt-guided localization reviewed below. C. Prompt-guided Localization Prompt-guided localization conditions a DoA estimator on a user-specified cue, enabling it to localize the queried tar- get selectively rather than all active sources indiscriminately. Existing work has explored this idea through two modalities: text-based conditioning for semantic descriptions and audio- based conditioning for acoustic enrollment. On the text side, early work already explored text-like cues in the form of fixed keywords. Keyword-based speaker localization [55] introduces the task of localizing the speaker 4 who uttered a predefined trigger phrase (e.g., a wake word) in the presence of overlapping speech. Beyond fixed keywords, class-conditioned localization provides a related formulation where the query is a target event class rather than free-form text. A class-conditioned SELD framework [56] utilizes one- hot class indicators as conditioning cues, injecting them into an ACCDoA-based baseline via FiLM layers. This design enables the model to selectively estimate the DoA of the queried class while treating concurrent events from other classes as interference during training. More recently, SEL [57] extends text conditioning to free-form captions and benchmarks fusion schemes that combine textual embeddings (e.g., CLAP, BERT, or FlanT5) with spatial audio features such as GCC-PHAT for selective DoA prediction and trajectory estimation. For audio-based localization, LocSelect [58] studies target speaker localization given an enrollment utterance of the same target speaker. It first predicts a speaker-dependent spectro- gram mask to suppress interferers and then estimates the target DoA from the filtered spectrogram using a long short- term memory (LSTM)-based localizer [58]. Alternatively, ap- proaches like GCC-Speaker [59] adapt classical spatial fea- tures directly, learning speaker-dependent weights for GCC- PHAT to enhance selectivity in multi-speaker scenarios. Overall, existing prompt-guided localization studies are predominantly unimodal, utilizing either text or audio cues but not both within a unified model. Moreover, evaluations are often confined to simplified settings, with limited coverage of moving targets, noisy mixtures, or prompts that correspond to multiple targets. I. TASK FORMULATION Let us consider a dual-channel recording setup. The signal received at the m-th microphone (m ∈ 1, 2), denoted by x m (t), is modeled as the sum of J source signals convolved with their respective room impulse responses (RIRs), together with additive ambient noise: x m (t) = J(t) X j=1 s j (t)∗ h m,j (t,θ j (t)) + n m (t),(1) where s j (t) is the j-th source signal, J (t) is the number of active sources at time t, n m (t) is the additive noise at the m- th microphone, and h m,j (t,θ j (t)) represents the RIR between the j-th source and the m-th microphone, which is critically dependent on the time-varying DoA of the source θ j (t). ⋆ indicates convolution. Our goal is to learn a prompt-guided localization model M ψ , parameterized by ψ, designed to estimate the spatial trajectory of a target source specified by a multimodal prompt. Unlike traditional SSL systems, the proposed SelectTSL is conditioned on a user-provided cue. It maps the dual-channel mixture signal x(t) = [x 1 (t),x 2 (t)] T , together with an aux- iliary audio cue x cue (t) or a textual prompt x text , to two corresponding frame-level predictions: • The DoA Posteriorgram, ˆ P DoA ∈ [0, 1] T out ×Θ denotes frame-level DoA confidence maps, where T out is the number of output frames after temporal alignment and Θ = 180 is the number of azimuth bins discretized from 0 ◦ to 179 ◦ at 1 ◦ resolution. Due to front-back ambiguity in symmetric dual-microphone arrays, predictions are restricted to a 180 ◦ range, and the multi-label formulation allows multiple active bins per frame with continuous confidence values. • The Source Cardinality, ˆ P card ∈ [0, 1] T out ×3 denotes frame-level cardinality predictions, a probability distribu- tion over the presence of 0, 1, or 2 active target sources. We optimize the model parameters ψ to jointly predict the target sources’ DoA posteriorgram ˆ P DoA and frame-level cardinality ˆ P card , as formally defined by: ( ˆ P DoA , ˆ P card ) =M ψ (x(t)|x cue (t),x text ).(2) IV. SELECTTSL ARCHITECTURE SelectTSL is an end-to-end framework for prompt-guided selective target sound localization, in which a PGSA mod- ule generates extraction-informed embeddings from a dual- channel mixture and a DoA estimator jointly predicts frame- level DoA and source cardinality. The overall architecture is shown in Fig. 2 and is described as follows. A. Prompt-Guided Selective Attention Module (PGSA) We introduce the PGSA module, which acts as a prompt- guided selective filter. It leverages multimodal prompts (text or audio) to extract target-specific representations from the input mixture. 1) Audio Encoder: Given a dual-channel signal x m (t) for m ∈ 1, 2, we compute complex spectrograms S m (τ,ν) by short-time Fourier transform (STFT) using a Hann window, where τ ∈ 1,...,T in and ν ∈ 1,...,F represent time frames and frequency bins, respectively. We extract • Magnitude cues: A m (τ,ν) = |S m (τ,ν)|, which are encoded by a Conv1D followed by a ReLU module into H (enc) m ∈R D emb ×T in : H (enc) m = ReLU(Conv1D(A m )).(3) • Spatial cues: IPD and inter-channel level difference (ILD): IPD(τ,ν) =∠S 1 (τ,ν)−∠S 2 (τ,ν),(4) ILD(τ,ν) = log |S 1 (τ,ν)| + ε |S 2 (τ,ν)| + ε .(5) 2) Prompt Encoders: We employ a frozen CLAP model to extract audio cue x cue (t) and text x text : c audio = MLP cue (CLAP audio (x cue (t)))∈R 256 ,(6) c text = MLP text (CLAP text (x text ))∈R 256 .(7) Here, MLP cue and MLP text serve as projection layers. During training, we freeze the CLAP encoders and optimize only the projection layers. We then form a unified guidance vector c fused = [c audio ;c text ] ∈R 512 . In our model, the audio prompt x cue (t) is taken as the first 1 s of a 6 s clip, and the remaining 5 s of that clip is used as the target segment for selection and localization. 3) Fusion Layer: We condition the encoded magnitude features H (enc) m using a two-stage FiLM cascade driven by c fused : FiLM(h,c) = γ(c fused )⊙ h + β(c fused ),(8) 5 Mixed Audio Audio Cue Audio Cue Encoder Text Encoder Fusion Layer Aligner Card Head Prompt-Guided Selective Attention Module (PGSA) DoA N sources =0, 1, 2, ... Angle Refiner FRBlocks TCNBlocks BiGRU Progressive Refinement Temporal Modeling Module (PRTM) DoA Estimator Text Cue Audio Encoder "Locate the cat" ❄ ❄ Projection & Concat ReLU& Conv ILD IPD C 0 &S 0 cos&sin Acoustic Branch IPD Enhancer Semantic Branch ILD IPD enh 0 90 180 DoA(°) time ✅ ❌ ❌ Norm FiLM Projection ×2 Fusion Layer (a) Fusion Layer and Extraction Network Chunking Intra-block Inter-block Cross-Attention Over-and-Add Extraction Network (b) IPD Enhancer TCN block d=1 TCN block d=2 TCN block d=4 TCN block d=8 Forward GRU Backward GRU DWConv1d Norm&ReLU SE Pre-proj PRTM (c) Progressive Refinement Temporal Modeling Module (PRTM) Ac head Sem head DWS Conv2d Gated-FiLM Fuse head & IPD Delta Acoustic Branch Semantic Branch IPD enh × Extraction Network Fig. 2. Overall architecture of SelectTSL. The PGSA module outputs the target magnitude spectrogram and a stack of EIEsH (eie) (from the Extraction Network). The DoA estimator enhances spatial cues with these EIEs and aggregates semantic & spatial information for localization and cardinality estimation. where γ(·) and β(·) are learnable functions that map the conditioning vector to scale and shift parameters. First, c fused modulates the full-dimensional H (enc) m , followed by a projec- tion. Then, a downsampled guidance modulates the projected features, yielding H (fused) m . 4) Extraction Network: We segment H (fused) m into overlap- ping chunks of length L chunk with 50% overlap, and process them using N dprnn dual-path blocks (intra-/inter-chunk RNNs from DPRNN [29]). Inside the stack, we periodically inject cross-attention between the current audio features and the semantic guidance c fused . Let H ∈R T×D h denote the frame- level features in a block and c fused ∈R D c the guidance vector. We broadcast the guidance across time as ̃ C ∈R T×D c and compute cross-attention: Q = ̃ CW Q ,K = HW K ,V = HW V ,(9) with W Q ∈R D c ×d k and W K ,W V ∈R D h ×d k . The attention weights and residual update are α = softmax QK ⊤ √ d k , H← H +αV.(10) After the DPRNN stack, we project its output to a soft mask M m , which filters the encoder features and a decoder reconstructs the target magnitude: ˆ H m = M m ⊙ H (enc) m , ˆ A m = MLP dec ( ˆ H m ).(11) We use the extracted target features ˆ A 1 ∈R F×T for DoA estimation, where F and T denote frequency bins and frames. From the last K EIE DPRNN blocks (before overlap-and- add), we export extraction-informed embeddings (EIEs). Each block’s features are linearly projected to a TF feature map in R F×T in , as illustrated in the Extraction Network (bottom-left) of Fig. 2. Averaging across the two channels yields H (eie) k ∈ R F×T in , k = 1,...,K EIE . Stacking forms H (eie) = stack H (eie) 1 ,...,H (eie) K EIE ∈R K EIE ×F×T in . (12) These maps emphasize TF regions that are relevant for extract- ing the target and serve as guidance for IPD enhancement. B. DoA Estimator The DoA estimator consists of two modules: an IPD En- hancer that refines spatial cues, and a Progressive Refinement Temporal Modeling Module (PRTM). Although the PGSA module outputs ˆ A m for each channel m, these per-channel magnitude features are highly redundant. Therefore, we use only ˆ A 1 as the magnitude input to the DoA estimator. The IPD Enhancer takes the mixture spatial cues (IPD, ILD), the extracted target features ˆ A 1 from the PGSA module, and the H (eie) from the Extraction Network. It then outputs the enhanced IPD IPD enh . The PRTM operates on ˆ A 1 , IPD enh , and ILD to produce (i) a frame-level DoA posteriorgram over Θ azimuth bins and (i) a categorical distribution over the source-count set N = 0, 1,...,N max (with N max = 2 in our experiments), where N max is the maximum number of simultaneously active target sources. 1) IPD Enhancer: We refine the mixture IPD with an IPD Enhancer that is explicitly conditioned on both low-level spatial cues and high-level EIEs (Fig. 2). Acoustic branch. The input to the acoustic branch com- prises ˆ A 1 , C 0 = cos(IPD), S 0 = sin(IPD), and ILD. These features are concatenated to form a 4-channel spectrogram 6 representation, which is subsequently processed by an acoustic feature encoder (denoted as ac head). This module employs a lightweight depthwise-separable Conv2d (DWS Conv2d) to efficiently capture fine-grained spectro-spatial features. Semantic branch. Operating in parallel, this branch ex- clusively processes H (eie) from the Extraction Network. It is encoded by a semantic head into a conditioning representation that drives a gated FiLM modulation. Specifically, global pooling is employed to derive channel-wise scale and shift parameters, while a concurrent spatial gate modulates the acoustic features to enable spatially selective IPD enhance- ment. Fusion and residual prediction. Let A ac denote the acous- tic features produced by the acoustic branch and A sem the semantic features produced by the semantic head. We fuse them by applying FiLM conditioning followed by a gated modulation: ̃ A = 1 + g(A sem ) ⊙ FiLM(A ac ,A sem ),(13) where g(·) is a spatial gating head that modulates the FiLM- conditioned features to enable spatially selective IPD en- hancement. A fuse head then predicts cosine–sine residuals (IPD deltas), and the enhanced IPD is recovered by the four- quadrant inverse tangent atan2(·,·): C 0 = cos(IPD), S 0 = sin(IPD), (∆C, ∆S) = ∆ IPD ( ̃ A), IPD enh = atan2 S 0 + ∆S, C 0 + ∆C ∈R F×T in . (14) 2) Progressive Refinement and Temporal Modeling (PRTM) Module: For each frame τ , we concatenate the features: z τ = ˆ A 1 (:,τ ); IPD enh (:,τ ); ILD(:,τ ) .(15) Stacking over time gives Z in ∈R T in ×C feat . The sequence is then processed by a PRTM module. First, the input is projected and processed by a stack of Feature Refinement Blocks (FRBs), whose composite operation is denoted as R(·). This module progressively fuses three distinct cues: target magni- tude, enhanced IPD, and ILD. Each FRB is implemented using depthwise-separable 1D convolutions, squeeze-and-excitation modules, and residual connections. This progressive fusion projects the inputs into a unified feature space where spatial and spectral evidence are tightly coupled. On top of this fused representation, a dilated temporal convolutional network (TCN) and a bidirectional gated recurrent unit (BiGRU) focus on modeling longer-range temporal patterns: H (temp) = BiGRU TCN R(Linear(Z in )) ∈R T in ×2D gru . (16) A temporal alignment layer adjusts the frame rate to obtain H (aligned) =T H (temp) ∈R T out ×2D gru .(17) The alignment layer T (·) compresses the sequence to a fixed length T out , so that the prediction heads operate at the same temporal resolution as the labels. 3) Prediction Heads: The aligned features H (aligned) ∈ R T out ×2D gru are decoded by the Angle Refiner (Fig. 2), which is implemented as two parallel heads. Let proj DoA (·) and proj card (·) denote two learned frame-wise affine projections that output logits, applied row-wise to H (aligned) : ˆ P DoA = σ proj DoA H (aligned) ∈ [0, 1] T out ×Θ ,(18a) ˆ P card = softmax proj card H (aligned) ∈ [0, 1] T out ×|N| , (18b) where proj DoA :R 2D gru →R Θ and proj card :R 2D gru →R |N| . During inference, we strictly couple the two predictions frame-wise. First, the number of active sources is estimated as ˆn t = arg max n∈N ˆ P card (t,n). Then, the DoA estimates are derived from the top-ˆn t peaks of ˆ P DoA (t, :). Specifically, peak picking involves: (i) circular Gaussian smoothing on ˆ P DoA (t, :) to mitigate jaggedness while handling the 0 ◦ -180 ◦ wrap-around; (i) identifying strict local maxima that exceed an adaptive threshold τ t = max(μ t + σ t ,τ min ) (where μ t ,σ t are the frame-wise mean and std, and τ min = 0.3); and (i) selecting the ˆn t highest peaks as the final DoA predictions. 4) Training Objective: The network is jointly trained with three loss terms. The separation loss is the negative SI-SNR averaged over the two channels: L sel =− 1 2 SI-SNR( ˆ s 1 ,s 1 ) + SI-SNR( ˆ s 2 ,s 2 ) .(19) For localization, we use a Binary Cross Entropy (BCE) for frame-level DoA estimation and a Cross Entropy (CE) for source-count prediction: L DoA = BCE ˆ P DoA ,P ⋆ DoA , L card = CE ˆ P card ,n ⋆ , (20) where P ⋆ DoA ∈ [0, 1] T out ×Θ is a soft DoA label built from ground-truth (GT) azimuths, and n ⋆ ∈ 0, 1, 2 T out is the frame-wise source-count label. The total loss is a weighted sum of separation, DoA and cardinality terms: L = ω sel L sel + ω DoA L DoA + ω card L card ,(21) where we fix ω sel = 1 and tune the other weights to balance the influence of each task’s gradient. This is crucial as the losses operate at different scales, especially near convergence. The selection loss L sel (negative SI-SNR) typically con- verges to a magnitude on the order of O(10 1 ). In contrast, L DoA is a BCE averaged over T out ×Θ bins. Due to the sparse nature of the DoA target (at most 2 of 180 bins are active), a converging model predicts near-zero loss for ≈ 99% of the outputs. Consequently, the average loss L DoA becomes very small, e.g., O(10 −1 ) or O(10 −2 ). To prevent the DoA gradient from vanishing and ensure its influence remains comparable to the separation term (i.e., ω DoA L DoA ≈ ω sel L sel ), we need to compensate for this scale mismatch. Therefore, ω DoA is chosen to be on the order of O(10 2 ). V. EXPERIMENTAL SETTING A. Dataset Our model is trained and evaluated on a synthesized dataset that covers a wide range of acoustic conditions. The synthesis process combines clean source signals with simulated dynamic RIRs and ambient noise. a) Sound Sources: To ensure diversity in our target sound events, we collected source signals from several public datasets. Specifically, speech signals were collected from the 7 TABLE I DATASET STATISTICS AND KEY SETTINGS. ItemValue Simulated dataset (tr/val/t) Hours: 288.9 / 31.1 / 18.1 Clips: 208,008 / 22,392 / 13,032 Real dataset (tr/val/t) TAU-SRIR: 9 rooms; Clips: 72,000 / 9,000 / 9,000 clips target categories397 Prompttext and/or 1 s audio cue Mixture1–2 targets/mixture; active/frame ∈0, 1, 2 Noisesignal-to-noise ratio (SNR) ∼U([−5, 5]) dB Clip format5 s, dual channel, 16 kHz Labeling Azimuth bins A=180 over [0, 180) Label rate 15 Hz (T out =75) LibriSpeech corpus [60]. Musical instrument sounds were collected from the C-Music Pianos [61] and GuitarSet datasets [62]. For a broader range of general acoustic events, we utilized samples from AudioSet [63] and WavCaps [64]. b) Noise Sources: A collection of noise recordings was used to create realistic and challenging noisy mixtures. These were sourced from multiple datasets, including MS- SNSD [65], WHAM! [66], ESC-50 [67], UrbanSound [68], QUT-NOISE [69] and Musan [70]. c) Data Simulation: We generated all training, valida- tion, and test mixtures synthetically. The process relies on dynamic RIRs generated using the GPURIR library [71]. For each simulated scenario, we define a room with dimensions of 4 × 4 × 2 meters and a reverberation time (T 60 ) of 0.2 seconds. A dual-channel microphone array is positioned at coordinates [1.9, 2.0, 1.0] and [2.1, 2.0, 1.0], with the inter- microphone distance of 20 cm. To simulate moving sources, we generate 5-s trajectories within a frontal azimuth range of 180 ◦ and uniformly resample them at 15 Hz, yielding T out = 75 discrete target directions per clip. Motivated by the DCASE SELD benchmarks, we use a label frame rate coarser than the acoustic feature frame rate to shorten sequences while retaining sufficient temporal detail [72]. At 15 Hz, a source traversing 180 ◦ in 5-s moves by at most 2.4 ◦ between consecutive labels, which is much smaller than the 20 ◦ angular tolerance adopted in the DCASE SELD localization metrics [8], [72]. Each mixture is generated with one or two moving sound sources. Since these sources may be intermittently silent, the number of active sources in any given time frame is zero, one, or two. These moving sources are created by convolving the clean source signals with their corresponding dynamic RIRs. Finally, a randomly selected noise signal is added to the mixture at SNR uniformly sampled from the range of [−5, 5] dB. In total, we generated 288.9 hours of audio for the training set, 31.1 hours for the validation set, and 18.1 hours for the test set. d) Real Recordings: We use a subset of nine rooms from TAU-SRIR [73]: Bomb shelter, Gym, PB132, PC226, SA203, SC203, SE203, TB103, and TC352. This subset is chosen to cover the main indoor archetypes represented in the dataset, span distinct surface materials and room volumes (rock and plastic-coated surfaces, carpet vs. hard floors, glass partitions), and include both trajectory types provided by the dataset (circular in Bomb shelter/Gym/PB132/PC226 and lin- ear in SA203/SC203/SE203/TB103/TC352). Although TAU- SRIR provides discrete source positions along circular/linear trajectories, we render each 5-s evaluation segment with a single spatial room impulse response (SRIR), i.e., a measured multichannel RIR with fixed source–microphone geometry, selecting the symmetric mic pair from the tetrahedral array to obtain two channels. The goal is to ensure diverse rever- beration and reflection patterns while keeping the evaluation compact and non-redundant. Clean 5-s utterances from Lib- riSpeech [60] are convolved with the selected two-channel RIRs (symmetric mic pair 0 and 2) and mixed with background noise at SNR uniformly sampled from [−5, 5] dB; angles are expressed in the two-microphone frame by subtracting the array baseline and folding to [0 ◦ , 180 ◦ ). B. Implementation details With a Hann STFT at 16 kHz (n fft = 1024, hop = 256) we obtain F = 513 and T in = 251; in the audio encoder, D emb = 256; in the PGSA module, K = 80, N dprnn = 6, L chunk = 128, D h = D c = d k = 64, K EIE = 2; in the DoA estimator, C feat = 2180, C tcn = 256, D gru = 256 (thus 2D gru = 512), T out = 75, Θ = 180, N max = 2. For training, we use up to 200 epochs with early stopping (patience = 10) and Adam optimization with lr = 5× 10 −4 , and gradient clipping with clip norm = 1. For the loss weights in the training objective, we use (ω sel ,ω DoA ,ω card ) = (1, 100, 1) in all experiments. C. Baseline Methods We categorize the baselines into four types: (1) Track-wise methods (IPDNet [37], EINV2 [51], embed-ACCDoA [74], SALSA-Lite [42], the DCASE 2025 Task 3 baseline (denoted as DCASE25 below) [52]), which allocate a fixed number of output tracks for DoA/activity trajectories; (2) SELD+DoA methods (SELDnet [7], SELDT [53]), which output per- class predictions without explicit track control; (3) Pure DoA methods (SRP-DNN [45], FN-SSL [36]), which estimate DoA from spatial features without activity branches; and (4) Prompt-based methods, where we compare the text-queried SEL model [57] against our proposed SelectTSL. For fair comparison, we adapt all baselines to our binaural two-channel setup while preserving their original prediction heads and training objectives as much as possible. Specifically, we adapt the baselines as follows: (i) multi-microphone or FOA inputs are reduced to the L-R pair, consistent with stereo/variable-array settings in prior work [37], [52], [75]; (i) standard FOA features are replaced by two-microphone spatial cues (cos(IPD), sin(IPD), ILD, and GCC-PHAT), following established feature substitution protocols used in DCASE baselines and SALSA-lite [42], [72], [76]; (i) the geometry is constrained to the horizontal plane (0 ◦ –180 ◦ azimuth) to re- move front-back ambiguity [52], [75]; (iv) original prediction heads and losses are preserved with minimal changes, in line with community practice and variable-array SSL designs [37], 8 [72] and (v) inference outputs are mapped to a common frame- wise DoA heatmap format for fair comparison [42], [77]. D. Metrics We report static frame-level metrics (MAE, Precision, F1, Recall) and dynamic trajectory-level metrics (MOTA ∗ , DetA, OSPA-T). The evaluation is based on framewise DoA matching. For each frame t, let P t = ˆ θ t,k denote the set of predicted azimuth hypotheses and G t = θ t,k the annotated references. We match P t to G t via a Hungarian assignment that minimizes the circular distance on [0, 180 ◦ ), i.e., d ◦ ( ˆ θ,θ) = min | ˆ θ − θ|, 180 ◦ − | ˆ θ − θ| . A match is a TP if its error ≤ δ; unmatched predictions/references are FP/FN. Let TP = P t TP t , FP = P t FP t , FN = P t FN t , and GT = P t |G t |. We report precision P = TP TP+FP , recall R = TP TP+FN and F1 = 2PR P +R . MAE (Mean Angular Error): MAE = 1 P t |M t | X t X ( ˆ θ,θ)∈M t d ◦ ( ˆ θ,θ), where M t denotes the Hungarian assignment between P t and G t in frame t. MOTA ∗ (Multiple Object Tracking Accuracy; ID-agnostic): We follow [54] but omit the ID–switch penalty; since all DoA hypotheses extracted under the same prompt belong to a single class, we report an ID-agnostic variant. MOTA ∗ = 1− P t (FN t + FP t ) P t |G t | = 1− FN + FP GT . DetA (Detection Accuracy): DetA = TP TP + FP + FN . OSPA-T (Optimal Sub-Pattern Assignment [78]). Let d c ( ˆ θ,θ) = mind ◦ ( ˆ θ,θ),c and n t = max(|P t |,|G t |). The frame score is OSPA t = 1 n t X ( ˆ θ,θ)∈M t d c ( ˆ θ,θ) + c(FP t + FN t ) , where the cutoff c in [78] is set to 15 ◦ , and FP t =|P t |−|M t |, FN t = |G t | − |M t |. The sequence score averages over all frames: OSPA-T = 1 T T X t=1 OSPA t . VI. EXPERIMENTS A. Main Results Table I reports the performance of all baselines across static frame-level metrics (MAE, Prec., F1, Recall) and dy- namic trajectory-level metrics (MOTA ∗ , DetA, OSPA-T). For readability, we discuss results by the four baseline families defined in Sec. V-C. To disentangle architectural effects from input conditioning, we report results under three input con- ditions: Mix, Sel-Joint, and Clean. Specifically, Mix denotes the original stereo mixture with an unknown number of active sources (0/1/2) plus noise, fed directly to the baseline without any extraction front-end. Sel-Joint denotes end-to-end training where the baseline is preceded by the same PGSA module as ours and takes the concatenation of the PGSA-enhanced magnitude and the original Mix phase as inputs, forming a two-channel representation. Clean is a clean condition where the input contains only the target source(s) consistent with the prompt, with no interfering sound sources and no noise. 1) Track-wise baselines: Within the track-wise family (IPDNet, EINV2, embed-ACCDoA, SALSA-Lite, DCASE25), only IPDNet and EINV2 remain effective in the Mix condition. IPDNet attains an MAE of 4.60 ◦ and an F1 of 0.24, whereas EINV2 achieves a higher F1 of 0.40 but at the cost of a much larger MAE of 11.37 ◦ . This apparent trade-off (high F1 but large MAE) is consistent with the DP-IPD and ACCDoA formulations. When the target is active, the per-track activity outputs are frequently triggered, leading to many detections and thus a higher F1. However, in the Mix condition the predicted DoAs are biased toward a direction between the target and interfering sources, which can still fall within the evaluation tolerance while increasing the average angular error. The remaining track-wise baselines have almost zero F1 and negative or near-zero MOTA ∗ on Mix, and only become viable once Sel-Joint or Clean inputs simplify the scene. Under Clean inputs, IPDNet and EINV2 achieve MAE ≈ 1.4 ◦ and MOTA ∗ ≈ 0.66–0.69, indicating that track- wise modeling can work once the prompt-based target is extracted. However, these systems are optimized to explain all active sources rather than follow a single target: DP-IPD based IPDNet treats any source at a given direction similarly, and ACCDoA-style systems (EINV2, SALSA-Lite, embed- ACCDoA, DCASE25) only weakly encode the prompt. In complex mixtures, tracks tend to follow non-target sources with similar geometry or break the target trajectory into short, fragmented segments. Our model directly targets the key weakness of generic track-wise designs: weak prompt binding that leads to drift and trajectory fragmentation in complex mixtures. We enforce target selectivity with a prompt-based PGSA front-end, whose EIEs condition the IPD Enhancer so that the refined spatial cues are biased toward the prompt- specified source. On top of these target-dominant cues, we predict a per-azimuth DoA posteriorgram together with a cardinality head, avoiding slot-based track competition and stabilizing tracking under time-varying activity, which yields much higher MOTA ∗ on Mix (Table I). 2) SELD+DoA (non-track-wise): The SELD+DoA base- lines (SELDnet, SELDT) predict frame-level activity and DoA per class without explicit tracks. In the Mix condition they are highly sensitive to interference: SELDnet attains MAE ≈ 15.8 ◦ , SELDT MAE ≈ 22.1 ◦ , and both have MOTA ∗ ≤ 0 with very poor detection. With Sel-Joint and Clean inputs, their static metrics improve substantially (MAE ≈ 6.6 ◦ ), consistent with prior SELD results on less cluttered scenes. Trajectory-level metrics improve much less. Even in Clean, SELDnet remains below MOTA ∗ = 0.12 and SELDT only reaches MOTA ∗ ≈ 0.19, suggesting that temporal inconsis- tency in the frame-level DoA predictions still leads to frag- mented trajectories. In comparison, SELDT is clearly stronger on trajectory-level metrics, confirming that its tracking- oriented objectives help stabilize trajectories. However, both models are trained to explain all event classes rather than a 9 TABLE I STATIC METRICS ARE FRAME-LEVEL (PEAK-MATCHING); DYNAMIC METRICS ARE TRAJECTORY-LEVEL. MAE IS THE MEAN ANGULAR ERROR ON TRUE POSITIVES. PREC. DENOTES PRECISION. MOTA ∗ IS ID-AGNOSTIC (NO ID-SWITCH PENALTY). OSPA-T IS REPORTED IN DEGREES WITH p=1 AND c=15. PREC./F1/RECALL/MOTA ∗ /DETA ARE REPORTED AS PERCENTAGES. † NON–TRACK-WISE: THE MODEL CANNOT SPECIFY THE NUMBER OF TRACKS WHEN PRODUCING DOA ESTIMATES. GLOBAL BEST RESULTS ARE IN BOLD; CATEGORY BEST RESULTS ARE UNDERLINED. CategoryModelInputStatic DoA (frame-level)Dynamic metrics (trajectory-level) MAE (°)↓ Prec. (%)↑ F1 (%)↑ Recall (%)↑ MOTA ∗ (%)↑ DetA (%)↑ OSPA-T (°)↓ Track-wise IPDNet Mix4.6074.6923.6315.6221.8827.0311.09 Sel-Joint4.1778.0042.6929.3921.1027.1410.23 Clean1.40 82.6281.0790.8369.3372.694.55 EINV2 Mix11.3747.2740.3631.167.9125.2811.26 Sel-Joint4.1869.1263.9949.5041.0452.407.66 Clean1.4576.2584.8095.5265.7773.624.15 embed-ACCDoA Mix20.0238.331.170.59-0.360.5914.74 Sel-Joint6.2859.9634.2023.937.9520.6310.86 Clean6.5160.0039.8329.809.9324.869.88 SALSA-Lite Mix19.6636.221.360.69-0.530.6814.72 Sel-Joint4.6670.1231.5320.3411.6718.7211.46 Clean4.3472.6037.0824.8915.5022.7610.78 DCASE25 Mix17.1543.340.690.35-0.110.3514.77 Sel-Joint5.7163.5735.0724.2110.3421.2610.79 Clean5.8064.4042.0631.2313.9726.639.72 SELD+DoA † SELDnet Mix15.8342.621.510.77-0.270.7614.71 Sel-Joint6.6257.6635.1825.316.7221.3410.54 Clean6.34 60.7641.3641.3611.1026.089.65 SELDT Mix22.1231.812.971.56-1.781.5114.58 Sel-Joint6.6661.5348.2639.7114.8831.818.40 Clean6.6362.96 53.1746.0118.9436.217.49 Pure DoA † SRP-DNN Mix6.2459.9131.3421.226.9818.5811.21 Sel-Joint5.56 63.8326.4916.717.2415.2711.99 Clean5.9762.8638.6327.8911.4123.9410.17 FN-SSL Mix7.0170.2317.4119.0818.2317.8710.09 Sel-Joint5.8382.5635.1222.9823.1025.137.82 Clean5.9279.7233.8224.2921.1323.788.84 Prompt-based SEL Mix5.3759.8230.4520.426.707.9613.14 Sel-Joint2.7883.0149.5044.1716.2523.717.96 Clean2.8482.9150.2343.8817.1421.178.01 OursMix0.9898.2895.6793.2091.5791.702.08 single prompt-based target. When multiple sources overlap, they distribute energy across several classes and directions instead of committing to a single, stable trajectory for the target. Consequently, even with Sel-Joint or Clean inputs their MOTA ∗ remains far below both the best track-wise baselines and our SelectTSL, which explicitly enforces selectivity and temporal consistency with respect to prompt. 3) Pure DoA (non-track-wise): The pure DoA baselines (SRP-DNN, FN-SSL) operate without any sound event de- tection (SED) head or explicit tracks, estimating DoA directly from binaural phase cues. SRP-DNN uses a causal CRNN to estimate a frame-wise spatial spectrum on a fixed azimuth grid, followed by peak-picking to obtain the DoA. In our binaural horizontal setup this yields very similar localization across input conditions (MAE = 6.24 ◦ / Mix, 5.56 ◦ / Sel- Joint, 5.97 ◦ / Clean) and only modest changes in MOTA ∗ (≈ 0.07/0.07/0.11). This limited variation reflects the fixed spatial grid and the absence of explicit temporal modeling: cleaner magnitudes sharpen the peaks but only mildly affect their centers, while framewise fluctuations still lead to frag- mented trajectories. FN-SSL also estimates DP-IPD but uses full-band / narrow- band fusion to exploit cross-frequency and temporal structure. From Mix to Sel-Joint and Clean, its MAE improves from 7.01 ◦ to 5.8 ◦ , and MOTA ∗ increases from 0.18 to 0.23 and 0.21, yielding a more favorable static–dynamic trade-off than SRP-DNN. Nevertheless, FN-SSL remains a purely geometric, single-stage DoA estimator that treats all spatial peaks in the mixture as equally plausible, without any mechanism to focus on the target source. Consequently, it significantly underperforms our SelectTSL integrating the PGSA front-end. This performance gap across both static and trajectory-level metrics confirms that geometry alone is insufficient for robust, prompt-based localization and tracking. 4) Prompt-based: The prompt-based baseline SEL [57] conditions localization directly on a text description of the tar- get event. It takes multichannel audio and a text query, fuses an audio representation with a CLAP-based text embedding, and predicts a discretized azimuth distribution per frame. Under the Mix condition, SEL already benefits from text conditioning 10 but remains limited by interference and ambiguous audio–text correspondence (MAE = 5.37 ◦ , MOTA ∗ = 0.07). Prepending the PGSA module (Sel-Joint) or switching to Clean inputs substantially strengthens frame-level localization (MAE ≈ 2.8 ◦ in both cases) and improves trajectory metrics (MOTA≈0.16–0.17). However, SEL predicts a single azimuth distribution per frame for a given text embedding. When more than one source matches the query, it often focuses on the dominant target or averages multiple directions, leading to missed or unstable secondary trajectories, consistent with the original benchmark. Our model is explicitly designed to handle up to at least two concurrent target sources per text query. Using the raw Mix as input, it achieves an MAE of 0.98 ◦ and a MOTA ∗ of 0.92, substantially outperforming SEL even when SEL is given Sel-Joint or Clean inputs. We attribute these gains to the combination of a prompt-based PGSA front-end and a per-azimuth DoA heatmap that represents multiple peaks per frame, rather than collapsing them into a single discrete angle. This yields much sharper and more stable trajectories, especially in frames where two target sources overlap. B. Ablation and Design Analysis of SelectTSL 1) System-level architectural ablations:We perform system-level ablations to quantify the contribution of key design choices in SelectTSL (Table IV), including: (i) coupling between selection and DoA (conditioning on extracted target magnitude vs. mixture magnitude), (i) spatial cues to the DoA estimator (IPD+ILD, w/o ILD, w/o IPD, or none), (i) the number K EIE of DPRNN blocks used to form EIEs, (iv) the DoA module design (with/without cross-attention and FuseLayer), (v) the cardinality head, and (vi) the training scheme (end-to-end vs. two-stage). In each row, only the factor in the “Setting” column is changed and all other components follow the main configuration. Across (i)–(i), both target-dependent conditioning and ex- plicit multichannel cues are indispensable. Even with EIEs disabled (K EIE =0), conditioning on the selected target clearly outperforms the Mix→DoA variant using mixture magnitude (MAE 1.21 ◦ vs. 3.32 ◦ , MOTA ∗ 0.88 vs. 0.49). Removing ILD already degrades performance, while removing IPD or all spatial cues leads to large errors (MAE ≥ 3.89 ◦ ) and much lower MOTA ∗ , confirming the central role of phase- based spatial information in reverberant multi-source scenes. For temporal and structural choices (i)-(vi), using EIEs from only a few DPRNN blocks is sufficient: our default K EIE =2 performs best, K EIE =1 remains close, whereas K EIE =3 or 4 monotonically worsen MAE and MOTA ∗ , suggesting that incorporating EIEs from more blocks mainly injects noise. Removing cross-attention in the Extraction Net- work hurts trajectory metrics more than replacing the Fu- sion Layer with simple concatenation, since cross-attention is where the prompt interacts with the mixture features. Re- moving it weakens target emphasis and increases interference, which hurts temporal association. The cardinality head is cru- cial for stable tracking: removing it sharply reduces MOTA ∗ (0.89→0.58) and increases OSPA (3.1 ◦ →6.3 ◦ ). Finally, the TABLE IV ABLATION STUDY OF SELECTTSL (MAE, F1, MOTA ∗ , OSPA-T;↓ LOWER IS BETTER,↑ HIGHER IS BETTER). F1 AND MOTA ∗ ARE REPORTED AS PERCENTAGES. THE FULL MODEL ROW IS SHOWN AS A REFERENCE AND IS NOT CONSIDERED WHEN HIGHLIGHTING BEST RESULTS. BEST RESULTS WITHIN EACH ABLATION GROUP ARE HIGHLIGHTED IN BOLD. SettingMAE (°)↓ F1 (%)↑ MOTA ∗ (%)↑ OSPA-T (°)↓ Full model (SelectTSL)0.9895.6791.572.08 Coupling: w/o EIEs (K EIE =0)1.2193.5887.932.72 Coupling: Mix→DoA3.3265.9249.167.82 Spatial: w/o ILD1.7987.1877.274.14 Spatial: w/o IPD3.8960.4343.308.39 Spatial: w/o spatial cues4.5454.5237.489.20 EIE blocks: 11.4391.6888.763.11 EIE blocks: 31.9584.9873.884.50 EIE blocks: 42.2581.7969.204.97 Backbone: w/o cross-attn1.5091.0483.553.24 Backbone: w/o FuseLayer1.0894.6589.842.40 Cardinality: w/o head2.9173.5758.196.33 Training: separate Sel/DoA1.2890.1681.643.56 two-stage training variant lags behind end-to-end optimization (MAE 1.28 ◦ , F1 0.90, MOTA ∗ 0.82), indicating that joint training is needed for PGSA to produce EIEs that are aligned with the DoA objective. 2) Prompt-only ablations: To disentangle the roles of text and audio cues at the fusion layer, we run prompt-only ablations where the mixture is always encoded but one con- ditioning branch is masked to an all-zero embedding, forcing the model to rely solely on the remaining modality (Table V). For text-only conditioning, we keep the text branch active and zero out the audio embedding. We use Qwen2.5-7B [79] to generate paraphrases for each class and group them into three bands by CLAP similarity to the canonical caption: high, medium, and low ([0.85, 1.00], [0.70, 0.85), [0.50, 0.70)). Performance degrades monotonically as similarity decreases: relative to the canonical “Text-Only” caption (MAE = 1.12 ◦ , MOTA ∗ = 0.86), the highest band already worsens to MAE = 1.94 ◦ /MOTA ∗ = 0.74, and the lowest band falls to MAE = 3.48 ◦ /MOTA ∗ = 0.43, indicating that semantically close prompts are crucial when no audio cue is available. For audio-only conditioning, we mask the text embedding and feed a 1 s audio cue through the audio-prompt branch. We compare four strategies: taking the first 1 s from the class clip (Audio-Only / Front), sampling a random 1 s segment (Pos- Rand), time-stretching the first 1 s (Aug-Front), and taking the first 1 s from another clip of the same class (XClip-Front). Using the clip onset yields the best audio-only performance (MAE = 1.57 ◦ , MOTA ∗ = 0.81); random or cross-clip cues substantially degrade it (MAE≈ 2.30 ◦ , MOTA ∗ ≈ 0.51), with simple time-stretch in between. This suggests that under audio- only conditioning, the model is most sensitive to temporal misalignment and identity mismatch, while moderate temporal deformation is less harmful. 3) DoA estimator ablations: We ablate the proposed DoA estimator to assess the IPD Enhancer, the semantic–acoustic branches, and the feature refinement block (FRB). Table VI summarizes MAE, frame-level F1, MOTA ∗ , and OSPA-T; below we focus on MAE and MOTA ∗ as representative static and trajectory metrics. 11 TABLE V JOINT ABLATION OVER TEXT AND AUDIO PROMPTS UNDER SINGLE-MODALITY CONDITIONING AT THE FUSION LAYER. F1 AND MOTA ∗ ARE REPORTED ON A 0–1 SCALE. ModalitySetting / Similarity MAE (°)↓ F1↑ MOTA ∗ ↑ OSPA-T (°)↓ Text-only Text-Only1.120.930.862.95 [0.85, 1.00]1.940.860.744.37 [0.70, 0.85)2.190.800.655.82 [0.50, 0.70)3.480.650.438.22 Audio-only Audio-Only1.570.890.813.67 Pos-Rand2.300.710.527.28 Aug-Front1.920.780.625.93 XClip-Front2.290.700.517.46 TABLE VI ABLATION ON DOA ESTIMATOR ARCHITECTURE. F1 AND MOTA ∗ ARE REPORTED ON A 0–1 SCALE. SettingMAE (°)↓ F1↑ MOTA ∗ ↑ OSPA-T (°)↓ Baseline and IPD / semantic–acoustic design (A1–A4) Full (ours)0.980.960.922.08 A1: w/o IPD Enhancer2.100.830.695.02 A2: semantic branch only1.400.910.843.32 A3: acoustic branch only1.620.890.803.78 A4: direct IPD input2.710.730.536.50 FRB design (B1–B7) B1: w/o FRB3.800.640.387.90 B2: FRB depth ×11.520.840.793.71 B3: FRB depth ×31.370.830.793.81 B4: FRB depth ×42.120.830.694.91 B5: FRB conv-only, ×2 (no SE)2.130.830.684.95 B6: FRB SE-only, ×2 (no conv)1.960.850.724.50 B7: FRB w/o residual, ×21.930.850.724.62 For the IPD Enhancer and semantic–acoustic design (A1–A4), removing the IPD Enhancer (A1) causes a marked degradation relative to the full model (MAE = 0.98 ◦ → 2.10 ◦ , MOTA ∗ = 0.92→ 0.69), showing that denoising and refining raw IPD is critical for robust localization. Using only the semantic branch (A2) or only the acoustic branch (A3) is also suboptimal: both fall behind the full system, and semantic-only outperforms acoustic-only, indicating that text-conditioned embeddings carry strong class cues but still benefit from being fused with spatial features. Replacing enhanced IPD with direct IPD input (A4) yields the worst performance in this group (MAE = 2.71 ◦ , MOTA ∗ = 0.53), further highlighting the importance of the IPD enhancement module. For the FRB variants (B1–B7), the results show that iterative refinement is necessary but must be carefully structured. Completely removing FRB (B1) severely hurts performance (MAE = 3.80 ◦ , MOTA ∗ = 0.38), indicating that a single pass through the backbone is insufficient. Varying FRB depth (B2–B4) suggests that the default depth ×2 strikes the best balance: shallower or deeper configurations do not improve MAE or MOTA ∗ and can even degrade them. Ablating internal components (B5–B7) by removing the convolution, SE branch, or residual connection again worsens both metrics, confirming that the full FRB with all three components is most effective. 4) Cardinality head ablations: As described in Sec- tion IV-B, we decouple DoA estimation and source-count pre- diction: the DoA head produces a frame-level posteriorgram over Θ azimuth bins, while a separate cardinality head predicts a distribution overN =0, 1, 2. At inference, we take ˆn t and select the top-ˆn t local maxima of ˆ P DoA (t, :) as DoA estimates (“Ours” in Table VII), so the cardinality head provides a discrete prior on the number of sources without explicitly TABLE VII ABLATION ON HOW THE CARDINALITY HEAD INTERACTS WITH THE DOA HEAD. ALL VARIANTS SHARE THE SAME BACKBONE AND CARDINALITY CLASSIFIER; ONLY THE WAY CARDINALITY INFORMATION IS USED DIFFERS. F1 AND MOTA ∗ ARE REPORTED ON A 0–1 SCALE. SettingMAE (°)↓ F1↑ MOTA ∗ ↑ OSPA-T (°)↓ Ours (top-ˆn peaks)0.980.960.922.08 Card-head attention1.870.850.724.67 Embed (concat to DoA head)1.520.910.833.45 injecting cardinality embeddings into the DoA representation. To test whether tighter coupling helps, we compare this design with two alternatives that feed cardinality information back into the DoA head. In the “Card-head attention” variant, the predicted count is mapped to an embedding that queries a multi-head attention block over DoA features, and the attention output is added via a residual connection. In the “Embed” variant, the 3-way cardinality probabilities are passed through a small MLP to produce an embedding that is concatenated with the DoA head input. Table VII shows that both variants are worse than the simple top-ˆn peak selection, despite using more parameters: MAE increases from 0.98 ◦ to 1.52 ◦ /1.87 ◦ and MOTA ∗ drops from 0.92 to 0.83/0.72. Compared with the “no cardinality head” configuration in Table IV, all three designs benefit from having a dedicated source-count predictor, but directly injecting car- dinality embeddings into DoA features tends to distort spatial structure. The decoupled design with top-ˆn selection therefore offers the best overall trade-off while keeping the architecture and decoding rule simple. C. Robustness and Generalization 1) Varying acoustic complexity: To study robustness under different acoustic conditions, we define three difficulty levels (A/B/C) by gradually enlarging the room and widening the range of reverberation time T 60 ; the exact parameter ranges are given in Table VIII. Level A uses compact rooms with short T 60 , Level B moderately increases both, and Level C spans the largest, most asymmetric rooms and the longest, most variable reverberation. As shown in Table VIII, performance degrades steadily from A to C. MAE increases from 2.36 ◦ to 3.03 ◦ and MOTA ∗ drops from 0.61 to 0.46, with similar trends for the other static and trajectory-level metrics. Larger, more asymmetric rooms introduce stronger spatial ambiguities, and longer, more vari- able T 60 smears binaural cues, together defining the robustness envelope of our model under increasing acoustic variability. 2) Varying motion dynamics: To evaluate robustness under different motion dynamics, we define four speed buckets (A– D) that control the instantaneous angular velocity of the sources. Bucket A corresponds to slow motion within ±5 ◦ , bucket B to moderate motion within ±15 ◦ , and buckets C and D to increasingly rapid motion within ±30 ◦ and ±50 ◦ , respectively (Table IX). Table IX shows that tracking performance degrades as motion becomes faster and more irregular. MAE remains low for slow and moderate motion (1.20 ◦ in A, 1.14 ◦ in B) but rises to 1.55 ◦ and 2.32 ◦ in buckets C and D, while MOTA ∗ drops from 0.96 to 0.53 with similar declines in the other 12 TABLE VIII PERFORMANCE AND ACOUSTIC CONFIGURATION ACROSS DIFFICULTY LEVELS. ROOM DIMENSIONS L x , L y , L z AND REVERBERATION TIME T 60 ARE SAMPLED FROM THE RANGES SHOWN FOR EACH LEVEL. SettingL x (m)L y (m)L z (m)T 60 (s)MAE ( ◦ )Prec.F1RecallMOTA*DetAOSPA-T ( ◦ ) Main4.04.02.00.200.980.980.960.930.920.922.08 A[3.6, 4.6] [3.6, 4.6] [1.8, 2.3] [0.20, 0.35]2.360.900.780.690.610.645.82 B[3.8, 6.0] [3.8, 5.2] [2.2, 3.0] [0.25, 0.55]2.650.880.720.610.520.566.82 C[3.5, 7.0] [3.5, 6.0] [2.2, 3.0] [0.20, 0.65]3.030.830.680.580.460.527.21 TABLE IX PERFORMANCE ACROSS DIFFERENT MOTION SPEED BUCKETS. Bucket Range ( ◦ ) MAE Prec.F1Recall MOTA* DetA OSPA-T ( ◦ ) A±51.200.980.980.980.960.961.63 B ±151.140.970.950.930.900.912.29 C ±301.550.950.880.820.780.783.98 D ±502.320.880.720.610.530.576.91 TABLE X ABLATION AND ROBUSTNESS UNDER DIFFERENT MOTION AND PROMPT CONDITIONS. F1 AND MOTA ∗ ARE REPORTED ON A 0–1 SCALE. CategorySettingMAE (°)↓ F1↑ MOTA ∗ ↑ OSPA-T (°)↓ Movement stat–stat0.150.990.980.45 mov–stat0.400.980.960.70 stat–mov0.970.960.922.02 mov–mov0.980.960.922.08 No-prompt r=30%1.000.960.921.20 r=50%2.220.830.673.32 r=70%2.530.780.604.55 metrics. Higher angular velocities induce larger frame-to- frame direction changes, breaking the temporal smoothness the model relies on, so gradual or moderate movement is handled well whereas abrupt high-speed rotations remain challenging. 3) Robustness to motion and prompt conditions: We further test robustness when spatial dynamics are present and when text prompts may be missing (Table X). For motion, we inde- pendently control movement of the background and the target, yielding four regimes: stat–stat (static noise, static source), mov–stat (moving noise, static source), stat–mov (static noise, moving source), and mov–mov (both moving), with all other factors fixed. For prompts, we create no-prompt clips where the caption is removed but the target remains in the mixture, and vary the no-prompt proportion r ∈30%, 50%, 70%. In the movement group, the fully static case (stat–stat) achieves the best performance (MAE = 0.15 ◦ , MOTA ∗ = 0.98), background motion alone has only a mild impact (mov– stat), while moving the target (stat–mov, mov–mov) increases MAE to about 0.97 ◦ and reduces MOTA ∗ to about 0.92. In the no-prompt group, performance degrades as the no-prompt proportion increases: MAE rises from 1.00 ◦ at r = 30% to 2.53 ◦ at r = 70%, while MOTA ∗ drops from 0.92 to 0.60. Even so, the model remains usable when captions are frequently absent, showing that it can still localize and track targets based primarily on audio cues. 4) Real-world data: We further assess generalization on the real-room subset of TAU-SRIR, using the same protocol as in simulation. As summarized in Table XI, the mean performance across rooms is MAE = 2.62 ◦ , MOTA ∗ = 0.77, and OSPA-T TABLE XI PERFORMANCE ON THE TAU-SRIR REAL-ROOM EVALUATION SET. SettingMAE ( ◦ ) Prec.F1Recall MOTA* DetAOSPA-T ( ◦ ) 01bombshelter2.820.860.890.920.770.802.92 02gym2.010.870.910.940.810.832.47 03pb1320.820.930.950.970.910.911.25 04pc2261.900.890.920.950.840.852.09 05sa2036.890.810.830.850.640.704.26 06sc2032.390.890.910.940.820.842.29 08se2030.790.790.810.830.610.684.74 09tb1033.370.850.880.910.750.783.04 10tc3522.630.870.900.930.790.822.62 Mean2.620.860.890.920.770.802.85 = 2.85 ◦ , indicating strong transfer to measured spaces. Room-level trends correlate with the measured room at- tributes and trajectory design. PB132 is a small carpeted classroom with circular trajectories, implying shorter effective reverberation and a stable direct-to-reverberant ratio (DRR); this yields easy association and correspondingly strong scores (MAE 0.82 ◦ , MOTA ∗ 0.91, OSPA-T 1.25 ◦ ). SA203 is a lecture hall with an inclined floor and linear trajectories at multiple ranges, which introduce stronger early/late reflections and larger DRR variation, leading to increased angular bias and track fragmentation (MAE = 6.89 ◦ , MOTA ∗ = 0.64, OSPA- T = 4.26 ◦ ). SE203 (a large classroom with hard floor and linear trajectories) shows a different failure mode: framewise azimuths remain sharp (MAE = 0.79 ◦ ), but repeated crossings and specular clutter cause association breaks, resulting in the highest OSPA-T (4.74 ◦ ). Overall, the model maintains accurate DoA estimates across diverse real rooms, with per- formance variations that can be explained by room geometry and trajectory complexity. VII. VISUALIZATION To better understand the effect of different prompt designs, we visualize the frame-wise DoA posteriors produced by four separately trained models with different prompt inputs: (i) a Full model that uses both the audio cue and the text prompt, (i) a Text-only model, (i) an Audio-only model that only uses the audio cue, and (iv) a No-prompt model. For each row in Fig. 3, all four models are evaluated on the same mixture recording, and the rightmost panel overlays their decoded DoA trajectories together with the GT. In the first row, only the Full model is able to follow the challenging non-stationary trajectory, while the Text-only and Audio-only models fail to provide a reliable track and the No-prompt model produces a completely wrong trajectory. In the second row, both the Full and Text-only models closely 13 Fig. 3. Heatmap visualization of frame-wise DoA posteriors from four separately trained models under different prompt configurations. For three representative target events (rows), the first four columns correspond to the Full (text+cue), Text-only, Audio-only, and No-prompt models, all evaluated on the same mixture. The rightmost column shows the estimated and GT DoA trajectories. Warmer colors indicate higher posterior probability. Fig. 4. Polar visualization of temporal-spatial DoA trajectories. The radial axis encodes time (0–75 frames) and the angular axis encodes DoA (0 ◦ -180 ◦ ). match the ground truth, whereas the Audio-only and No- prompt models exhibit large deviations and fragmented tracks. In the third row, the Full and Audio-only models roughly capture the motion pattern, while the Text-only and No-prompt models again fail to track the source. These qualitative results demonstrate that text and audio prompts provide complemen- tary information, and that jointly exploiting both leads to the most robust localization performance. We also visualize estimated DoA trajectories in polar co- ordinates in Fig. 4, where the radial axis encodes time (0–75 frames) and the angular axis encodes DoA (0 ◦ -180 ◦ ). Panel (a) shows a dual-source recording with wide angular motion. The model follows the target across the full range. Panel (b) presents a single-source sequence with smooth motion, where predictions form a continuous track well aligned with the ground truth. Panel (c) shows a two-source sequence with silent periods that create three disjoint active segments. The model quickly re-locks onto the correct direction whenever the target becomes active again. These examples illustrate that the model can handle dual-source, single-source, and silent-period cases while maintaining temporally consistent trajectories. VIII. CONCLUSION We presented SelectTSL, a framework for prompt-guided selective target sound localization that leverages prompt- based selectivity to learn target-aware representations for dual- channel DoA estimation. By leveraging text and audio prompts and a target-aware multi-spatial-cue representation, our model can selectively estimate the DoA of the target source. Exper- iments show consistent improvements on both static frame- level and dynamic trajectory-level metrics over competitive baselines. Future work will extend SelectTSL toward unified multimodal prompting (e.g., incorporating visual prompts and scene context) to facilitate robust deployment in real-world environments. REFERENCES [1] J. Benesty, J. Chen, and Y. Huang, Microphone array signal processing. Springer, 2008. 14 [2] R. H ̈ eb-Umbach, T. Nakatani, M. Delcroix, C. Boeddeker, and T. Ochiai, “Microphone array signal processing and deep learning for speech enhancement: Combining model-based and data-driven approaches to parameter estimation and filtering [special issue on model-based and data-driven audio signal processing],” IEEE Signal Processing Maga- zine, vol. 41, no. 6, p. 12–23, 2025. [3] C. Knapp and G. Carter, “The generalized correlation method for estimation of time delay,” IEEE Trans. on Audio, Speech, and Language Processing, vol. 24, no. 4, p. 320–327, 2003. [4] R. Schmidt, “Multiple emitter location and signal parameter estimation,” IEEE transactions on antennas and propagation, vol. 34, no. 3, p. 276– 280, 1986. [5] S. Chakrabarty and E. A. Habets, “Multi-speaker doa estimation using deep convolutional networks trained with noise signals,” IEEE J. Sel. Topics Signal Process., vol. 13, no. 1, p. 8–21, 2019. [6] E. C. Cherry, “Some experiments on the recognition of speech, with one and with two ears,” J. Acoust. Soc. Am., vol. 25, no. 5, p. 975–979, 1953. [7] S. Adavanne, A. Politis, J. Nikunen, and T. Virtanen, “Sound event localization and detection of overlapping sources using convolutional recurrent neural networks,” IEEE J. Sel. Topics Signal Process., vol. 13, no. 1, p. 34–48, 2018. [8] A. Politis, A. Mesaros, S. Adavanne, T. Heittola, and T. Virtanen, “Overview and evaluation of sound event localization and detection in dcase 2019,” IEEE Trans. on Audio, Speech, and Language Processing, vol. 29, p. 684–698, 2020. [9] A. Senocak, H. Ryu, J. Kim, T.-H. Oh, H. Pfister, and J. S. Chung, “Toward interactive sound source localization: Better align sight and sound!” IEEE Trans. Pattern Anal. Mach. Intell., 2025. [10] Q. Wang, H. Muckenhirn, K. Wilson, P. Sridhar, Z. Wu, J. R. Hershey, R. A. Saurous, R. J. Weiss, Y. Jia, and I. L. Moreno, “VoiceFilter: Targeted voice separation by speaker-conditioned spectrogram masking,” INTERSPEECH, p. 2728–2732, 2019. [11] X. Liu, H. Liu, Q. Kong, X. Mei, J. Zhao, Q. Huang, M. D. Plumbley, and W. Wang, “Separate what you describe: Language-queried audio source separation,” in INTERSPEECH, 2022, p. 1801–1805. [12] K. Saijo, J. Ebbers, F. G. Germain, S. Khurana, G. Wiehern, and J. Le Roux, “Leveraging audio-only data for text-queried target sound extraction,” in ICASSP. IEEE, 2025, p. 1–5. [13] M. Kim, R. Mira, H. Chen, S. Petridis, and M. Pantic, “Contextual speech extraction: Leveraging textual history as an implicit cue for target speech extraction,” in ICASSP. IEEE, 2025, p. 1–5. [14] T. Afouras, J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Deep audio-visual speech recognition,” IEEE Trans. Pattern Anal. Mach. Intell., 2018. [15] K. Li, F. Xie, H. Chen, K. Yuan, and X. Hu, “An audio-visual speech separation model inspired by cortico-thalamo-cortical circuits,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 10, p. 6637–6651, 2024. [16] Z. Pan, M. Ge, and H. Li, “Usev: Universal speaker extraction with visual cue,” IEEE Trans. on Audio, Speech, and Language Processing, vol. 30, p. 3032–3045, 2022. [17] R. Tao, X. Qian, Y. Jiang, J. Li, J. Wang, and H. Li, “Audio-visual target speaker extraction with reverse selective auditory attention,” IEEE Trans. on Audio, Speech, and Language Processing, 2025. [18] R. Gu, L. Chen, S.-X. Zhang, J. Zheng, Y. Xu, M. Yu, D. Su, Y. Zou, and D. Yu, “Neural spatial filter: Target speaker speech separation assisted with directional information.” in INTERSPEECH, 2019, p. 4290–4294. [19] R. Gu, S.-X. Zhang, Y. Zou, and D. Yu, “Towards unified all-neural beamforming for time and frequency domain speech separation,” IEEE Trans. on Audio, Speech, and Language Processing, vol. 31, p. 849– 862, 2022. [20] K. ˇ Zmol ́ ıkov ́ a, M. Delcroix, K. Kinoshita, T. Ochiai, T. Nakatani, L. Bur- get, and J. ˇ Cernock ` y, “Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,” IEEE J. Sel. Topics Signal Process., vol. 13, no. 4, p. 800–814, 2019. [21] C. Xu, W. Rao, E. S. Chng, and H. Li, “Spex: Multi-scale time domain speaker extraction network,” IEEE Trans. on Audio, Speech, and Language Processing, vol. 28, p. 1370–1384, 2020. [22] K. Liu, Z. Du, X. Wan, and H. Zhou, “X-sepformer: End-to-end speaker extraction network with explicit optimization on speaker confusion,” in ICASSP. IEEE, 2023, p. 1–5. [23] H.-W. Dong, N. Takahashi, Y. Mitsufuji, J. McAuley, and T. Berg- Kirkpatrick, “Clipsep: Learning text-queried sound separation with noisy unlabeled videos,” arXiv preprint arXiv:2212.07065, 2022. [24] Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in ICASSP.IEEE, 2023, p. 1–5. [25] H. Ma, Z. Peng, X. Li, M. Shao, X. Wu, and J. Liu, “Clapsep: Lever- aging contrastive pre-trained model for multi-modal query-conditioned target sound extraction,” IEEE Trans. on Audio, Speech, and Language Processing, 2024. [26] X. Liu, Q. Kong, Y. Zhao, H. Liu, Y. Yuan, Y. Liu, R. Xia, Y. Wang, M. D. Plumbley, and W. Wang, “Separate anything you describe,” IEEE Trans. on Audio, Speech, and Language Processing, 2024. [27] K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “Zero-shot audio source separation through query-based learning from weakly-labeled data,” in AAAI, vol. 36, no. 4, 2022, p. 4441–4449. [28] Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time– frequency magnitude masking for speech separation,” IEEE Trans. on Audio, Speech, and Language Processing, vol. 27, no. 8, p. 1256–1266, 2019. [29] Y. Luo, Z. Chen, and T. Yoshioka, “Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation,” in ICASSP. IEEE, 2020, p. 46–50. [30] Z.-Q. Wang, S. Cornell, S. Choi, Y. Lee, B.-Y. Kim, and S. Watanabe, “Tf-gridnet: Integrating full-and sub-band modeling for speech separa- tion,” IEEE Trans. on Audio, Speech, and Language Processing, vol. 31, p. 3221–3236, 2023. [31] P.-A. Grumiaux, S. Kiti ́ c, L. Girin, and A. Gu ́ erin, “A survey of sound source localization with deep learning methods,” J. Acoust. Soc. Am., vol. 152, no. 1, p. 107–151, 2022. [32] A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. S. Kweon, “Learning to localize sound sources in visual scenes: Analysis and applications,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 43, no. 5, p. 1605–1619, 2019. [33] H. Xuan, Z. Wu, J. Yang, B. Jiang, L. Luo, X. Alameda-Pineda, and Y. Yan, “Robust audio-visual contrastive learning for proposal-based self-supervised sound source localization in videos,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 7, p. 4896–4907, 2024. [34] Z. Song, J. Zhang, Y. Wang, J. Fan, and Z. Zhang, “Enhancing sound source localization via false negative elimination,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 12, p. 10 499–10 514, 2024. [35] T. N. T. Nguyen, K. N. Watcharasupat, N. K. Nguyen, D. L. Jones, and W.-S. Gan, “Salsa: Spatial cue-augmented log-spectrogram features for polyphonic sound event localization and detection,” IEEE Trans. on Audio, Speech, and Language Processing, vol. 30, p. 1749–1762, 2022. [36] Y. Wang, B. Yang, and X. Li, “Fn-ssl: Full-band and narrow-band fusion for sound source localization,” in INTERSPEECH, 2023, p. 3779–3783. [37] —, “Ipdnet: A universal direct-path ipd estimation network for sound source localization,” IEEE Trans. on Audio, Speech, and Language Processing, 2024. [38] P.-A. Grumiaux, S. Kiti ́ c, P. Srivastava, L. Girin, and A. Gu ́ erin, “Salad- net: Self-attentive multisource localization in the ambisonics domain,” in IEEE Workshop Appl. Signal Process. Audio Acoust.IEEE, 2021, p. 336–340. [39] W. Huang, Q. Huang, L. Ma, and C. Wang, “Swg-former: A sliding- window graph convolutional network for simultaneous spatial-temporal information extraction in sound event localization and detection,” arXiv preprint arXiv:2310.14016, 2023. [40] J. Hu, Y. Cao, M. Wu, Q. Kong, F. Yang, M. D. Plumbley, and J. Yang, “Sound event localization and detection for real spatial sound scenes: Event-independent network and data augmentation chains,” arXiv preprint arXiv:2209.01802, 2022. [41] Y. Shul and J.-W. Choi, “Cst-former: Transformer with channel-spectro- temporal attention for sound event localization and detection,” in ICASSP. IEEE, 2024, p. 8686–8690. [42] T. N. T. Nguyen, D. L. Jones, K. N. Watcharasupat, H. Phan, and W.-S. Gan, “Salsa-lite: A fast and effective feature for polyphonic sound event localization and detection with microphone arrays,” in ICASSP. IEEE, 2022, p. 716–720. [43] A. Berg, J. Engman, J. Gulin, K. Astr ̈ om, M. Oskarsson, and B. Sony Eu- rope, “The lu system for dcase 2024 sound event localization and detection challenge,” DCASE2024 Challenge, Tech. Rep, Tech. Rep., 2024. [44] A. Berg, J. Engman, J. Gulin, K. ̊ Astr ̈ om, and M. Oskarsson, “Learning multi-target tdoa features for sound event localization and detection,” arXiv preprint arXiv:2408.17166, 2024. [45] B. Yang, H. Liu, and X. Li, “Srp-dnn: Learning direct-path phase difference for multiple moving sound source localization,” in ICASSP. IEEE, 2022, p. 721–725. [46] H. Yin, M. Ge, Y. Fu, G. Zhang, L. Wang, L. Zhang, L. Qiu, and J. Dang, “Mimo-doanet: Multi-channel input and multiple outputs doa 15 network with unknown number of sound sources,” arXiv preprint arXiv:2207.07307, 2022. [47] K. Shimada, Y. Koyama, N. Takahashi, S. Takahashi, and Y. Mitsufuji, “Accdoa: Activity-coupled cartesian direction of arrival representation for sound event localization and detection,” in ICASSP.IEEE, 2021, p. 915–919. [48] K. Shimada, Y. Koyama, S. Takahashi, N. Takahashi, E. Tsunoo, and Y. Mitsufuji, “Multi-accdoa: Localizing and detecting overlapping sounds from the same class with auxiliary duplicating permutation invariant training,” in ICASSP. IEEE, 2022, p. 316–320. [49] Y. Cao, T. Iqbal, Q. Kong, Y. Zhong, W. Wang, and M. D. Plumbley, “Event-independent network for polyphonic sound event localization and detection,” arXiv preprint arXiv:2010.00140, 2020. [50] Y. Cao, T. Iqbal, Q. Kong, F. An, W. Wang, and M. D. Plumbley, “An improved event-independent network for polyphonic sound event localization and detection,” in ICASSP. IEEE, 2021, p. 885–889. [51] D. Mu, Z. Zhang, and H. Yue, “Mff-einv2: Multi-scale feature fusion across spectral-spatial-temporal domains for sound event localization and detection,” arXiv preprint arXiv:2406.08771, 2024. [52] K. Shimada, A. Politis, I. R. Roman, P. Sudarsanam, D. Diaz-Guerra, R. Pandey, K. Uchida, Y. Koyama, N. Takahashi, T. Shibuya et al., “Stereo sound event localization and detection with onscreen/offscreen classification,” arXiv preprint arXiv:2507.12042, 2025. [53] S. Adavanne, A. Politis, and T. Virtanen, “Localization, detection and tracking of multiple moving sound sources with a convolutional recurrent neural network,” arXiv preprint arXiv:1904.12769, 2019. [54] —, “Differentiable tracking-based training of deep learning sound source localizers,” in IEEE Workshop Appl. Signal Process. Audio Acoust. IEEE, 2021, p. 211–215. [55] S. Sivasankaran, E. Vincent, and D. Fohr, “Keyword-based speaker localization: Localizing a target speaker in a multi-speaker environment,” in INTERSPEECH, 2018. [56] O. Slizovskaia, G. Wichern, Z.-Q. Wang, and J. Le Roux, “Locate this, not that: Class-conditioned sound event doa estimation,” in ICASSP. IEEE, 2022, p. 711–715. [57] J. Zhao, X. Qian, Y. Xu, H. Liu, Y. Cao, D. Berghi, and W. Wang, “Text-queried target sound event localization,” in EUSIPCO.IEEE, 2024, p. 261–265. [58] Y. Chen, X. Qian, Z. Pan, K. Chen, and H. Li, “Locselect: Target speaker localization with an auditory selective hearing mechanism,” in ICASSP. IEEE, 2024, p. 8696–8700. [59] G. Li, W. Xue, W. Liu, J. Yi, and J. Tao, “Gcc-speaker: Target speaker localization with optimal speaker-dependent weighting in multi-speaker scenarios,” in ICASSP. IEEE, 2023, p. 1–5. [60] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in ICASSP.IEEE, 2015, p. 5206–5210. [61] M. Zhou, S. Wu, S. Ji, Z. Li, and W. Li, “A holistic evaluation of piano sound quality,” in National Conf. on Sound and Music Technology. Springer, 2023, p. 3–17. [62] Q. Xi, R. M. Bittner, J. Pauwels, X. Ye, and J. P. Bello, “Guitarset: A dataset for guitar transcription.” in ISMIR, 2018, p. 453–460. [63] J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in ICASSP. IEEE, 2017, p. 776–780. [64] X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang, “Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,” IEEE Trans. on Audio, Speech, and Language Processing, vol. 32, p. 3339– 3354, 2024. [65] C. K. Reddy, V. Gopal, R. Cutler, E. Beyrami, R. Cheng, H. Dubey, S. Matusevych, R. Aichner, A. Aazami, S. Braun et al., “The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,” arXiv preprint arXiv:2005.13981, 2020. [66] G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. L. Roux, “Wham!: Extending speech separation to noisy environments,” arXiv preprint arXiv:1907.01160, 2019. [67] K. J. Piczak, “ESC: Dataset for Environmental Sound Classification,” in ACM Int. Conf. Multimedia. ACM Press, p. 1015–1018. [68] J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in ACM Int. Conf. Multimedia, Orlando, FL, USA, Nov. 2014, p. 1041–1044. [69] D. Dean, S. Sridharan, R. Vogt, and M. Mason, “The qut-noise- timit corpus for evaluation of voice activity detection algorithms,” in INTERSPEECH.International Speech Communication Association, 2010, p. 3110–3113. [70] D. Snyder, G. Chen, and D. Povey, “Musan: A music, speech, and noise corpus,” arXiv preprint arXiv:1510.08484, 2015. [71] D. Diaz-Guerra, A. Miguel, and J. R. Beltran, “gpurir: A python library for room impulse response simulation with gpu acceleration,” Multimedia Tools and Applications, vol. 80, no. 4, p. 5653–5671, 2021. [72] A. Politis, S. Adavanne, and T. Virtanen, “A dataset of reverberant spatial sound scenes with moving sources for sound event localization and detection,” arXiv preprint arXiv:2006.01919, 2020. [73] A. Politis, S. Adavanne, T. Virtanen, E. Fagerlund, A. Koskimies, A. Hakala, and A. Gohar, “TAU Spatial Room Impulse Response Database (TAU-SRIR DB),” Apr. 2022. [74] K. Shimada, K. Uchida, Y. Koyama, T. Shibuya, S. Takahashi, Y. Mit- sufuji, and T. Kawahara, “Zero-and few-shot sound event localization and detection,” in ICASSP. IEEE, 2024, p. 636–640. [75] J. Wilkins, M. Fuentes, L. Bondi, S. Ghaffarzadegan, A. Abavisani, and J. P. Bello, “Two vs. four-channel sound event localization and detection,” arXiv preprint arXiv:2309.13343, 2023. [76] Y. Zhang, S. Wang, Z. Li, K. Guo, S. Chen, and Y. Pang, “Data augmentation and class-based ensembled cnn-conformer networks for sound event localization and detection,” Proc. DCASE, vol. 2021, 2021. [77] A. Politis, K. Shimada, P. Sudarsanam, S. Adavanne, D. Krause, Y. Koyama, N. Takahashi, S. Takahashi, Y. Mitsufuji, and T. Virtanen, “Starss22: A dataset of spatial recordings of real scenes with spatiotem- poral annotations of sound events,” arXiv preprint arXiv:2206.01948, 2022. [78] ́ A. F. Garc ́ ıa-Fern ́ andez, A. S. Rahmathullah, and L. Svensson, “A metric on the space of finite sets of trajectories for evaluation of multi-target tracking algorithms,” IEEE Trans. on Signal Processing, vol. 68, p. 3917–3928, 2020. [79] A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, G. Dong, H. Wei, H. Lin, J. Tang, J. Wang, J. Yang, J. Tu, J. Zhang, J. Ma, J. Xu, J. Zhou, J. Bai, J. He, J. Lin, K. Dang, K. Lu, K. Chen, K. Yang, M. Li, M. Xue, N. Ni, P. Zhang, P. Wang, R. Peng, R. Men, R. Gao, R. Lin, S. Wang, S. Bai, S. Tan, T. Zhu, T. Li, T. Liu, W. Ge, X. Deng, X. Zhou, X. Ren, X. Zhang, X. Wei, X. Ren, Y. Fan, Y. Yao, Y. Zhang, Y. Wan, Y. Chu, Y. Liu, Z. Cui, Z. Zhang, and Z. Fan, “Qwen2 technical report,” arXiv preprint arXiv:2407.10671, 2024.