Paper deep dive
A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers
Inho Kim, Sumyeong Ahn
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/28/2026, 3:10:36 AM
Summary
The paper introduces STeReO (Speech and Text Reranking Orchestrator), a cross-modal reranker designed to aggregate and rank heterogeneous candidates from speech and text retrievers for Retrieval-Augmented Generation (RAG) systems. To address the lack of training data, the authors curate a novel dataset using a foundation Audio Language Model (ALM) to provide relevance rankings. STeReO is trained using Low-Rank Adaptation (LoRA) with pointwise, pairwise, and listwise objectives. Experiments on Spoken SQuAD and MS MARCO demonstrate that STeReO significantly improves downstream question-answering performance compared to single-modality baselines and simple score fusion methods.
Entities (11)
Relation Signals (9)
STeReO → evaluatedon → MS MARCO
confidence 95% · and MS MARCO [2], comprising text-only web passages.
STeReO → evaluatedon → Spoken SQuAD
confidence 95% · We evaluate STeReO using a fixed top-k pipeline (k=5) on two distinct datasets: Spoken SQuAD [15]
STeReO → usestechnique → Low-Rank Adaptation
confidence 95% · It is fine-tuned using Low-Rank Adaptation (LoRA) and operates via a unified token-based scoring mechanism.
GPT-4o-Audio-Preview → usedfor → Dataset Annotation
confidence 93% · we utilize a foundation ALM (e.g., gpt-4o-audio-preview) as a unified evaluator... to establish a gold-standard ranking
STeReO → improves → Retrieval-Augmented Generation
confidence 92% · Our results demonstrate that the proposed algorithm excels at selecting the most relevant evidence, thereby significantly improving downstream question-answering performance.
STeReO → backbone → Qwen-Audio-Chat
confidence 90% · Our experiments evaluate three audio-native backbone architectures as student rerankers: ... Qwen-Audio-Chat [8]...
STeReO → backbone → Ultravox
confidence 90% · Our experiments evaluate three audio-native backbone architectures as student rerankers: Ultravox [11]...
STeReO → backbone →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Retrieval-Augmented Generation (RAG) systems have attracted significant interest for their ability to mitigate hallucinations in Large Language Models (LLMs). Although knowledge databases for RAG are increasingly diversifying to include various modalities such as speech and text, research on handling such multi-modal database scenarios remains limited. In this paper, we propose STeReO (Speech and Text Reranking Orchestrator), a reranker based on speech and text retrievers that aggregates disparate modality databases. To address the lack of specialized training data, we first curate a dataset comprising queries, mixed-modality evidence, and their corresponding relevance ranks. We then train the reranker and evaluate its effectiveness in both single-modality and mixed-modality scenarios. Our results demonstrate that the proposed algorithm excels at selecting the most relevant evidence, thereby significantly improving downstream question-answering performance.
Tags
Links
- Source: https://arxiv.org/abs/2608.26194v1
- Canonical: https://arxiv.org/abs/2608.26194v1
Trouble viewing inline? Open PDF directly →
Full Text
32,069 characters extracted from source content.
Expand or collapse full text
Kim Ahn A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers Inho Sumyeong Abstract Retrieval-Augmented Generation (RAG) systems have attracted significant interest for their ability to mitigate hallucinations in Large Language Models (LLMs). Although knowledge databases for RAG are increasingly diversifying to include various modalities such as speech and text, research on handling such multi-modal database scenarios remains limited. In this paper, we propose STeReO (Speech and Text Reranking Orchestrator), a reranker based on speech and text retrievers that aggregates disparate modality databases. To address the lack of specialized training data, we first curate a dataset comprising queries, mixed-modality evidence, and their corresponding relevance ranks. We then train the reranker and evaluate its effectiveness in both single-modality and mixed-modality scenarios. Our results demonstrate that the proposed algorithm excels at selecting the most relevant evidence, thereby significantly improving downstream question-answering performance. keywordsmultimodal reranking, heterogeneous candidates, speech retrieval †address: 1 Korea Institute of Energy Technology, South Korea †email: inho20@kentech.ac.kr, sumyeongahn@kentech.ac.kr 1 Introduction Retrieval-Augmented Generation (RAG) [14] enhances Large Language Models (LLMs) [1, 27, 7] by incorporating external knowledge, thereby mitigating hallucination [13] and grounding responses in external evidence. While most existing RAG pipelines assume text-only knowledge bases (Figure 1(a)), there is a growing need to integrate unstructured spoken content such as lectures and meeting recordings. However, this integration is non-trivial. The most straightforward approach is leveraging Automatic Speech Recognition (ASR), which converts audio to text before applying standard text retrieval. However, ASR introduces transcription errors that propagate through the pipeline, adds latency, and discards non-verbal information present in the original audio. Consequently, extending RAG to natively support spoken modalities is a significant, yet largely unexplored, challenge. Figure 1: Comparison of three RAG systems: (a) Text-based, (b) Speech-based, and (c) Text+Speech-based RAG. Several studies address ASR-related limitations through ASR-free speech retrieval methods. For instance, VoxRAG [22] employs direct speech-to-speech matching without ASR, while SpeechRAG [18] aligns textual queries with speech embeddings. However, their retrieval spaces are still restricted to single-modality, speech-only corpora (Figure 1(b)). In contrast, WavRAG [6] facilitates heterogeneous retrieval by projecting independent audio and text databases into a shared embedding space. Despite this advancement, such joint embedding approaches for multimodal bases suffer from the well-known modality gap [16, 10]. Even in audio-text models like CLAP [9], systematic score imbalance [23] often leads one modality to dominate retrieval results, regardless of its actual relevance. An alternative to avoid this gap is to employ modality-specific retrievers independently and merge their candidates via late fusion (Figure 1(c)). In this paradigm, a robust cross-modal reranker is essential for accurately evaluating and prioritizing candidates retrieved from diverse modalities. While recent listwise reranking methods, such as those based on permutation-invariant cross-encoders [24] or Fusion-in-Decoder architectures [30], have shown promising results on text-only pools, they remain inherently unimodal. Consequently, they fail to facilitate the complex cross-modal comparison necessitated by heterogeneous candidate pools. A primary hurdle in developing such a cross-modal reranker is the absence of training data: existing retrieval datasets lack explicit cross-modal relevance judgments required to directly compare audio and text candidates. To overcome this challenge, we first construct a novel dataset that explicitly captures the cross-modal relevance rankings among candidates generated by heterogeneous retrievers. Leveraging this dataset, we propose STeReO (Speech and Text Reranking Orchestrator), a cross-modal reranker designed to systematically align and prioritize heterogeneous candidates –drawn from independent, modality-specific databases– into a single, unified ranked list for a given text query. In summary, our main contributions are as follows: • First, we construct a novel cross-modal dataset that provides explicit relevance rankings for candidates retrieved from heterogeneous modalities. To achieve this efficiently, we propose a novel score fusion method that merges candidate sets from disjoint retrievers. This approach effectively filters the vast evidence space, allowing us to accurately extract relevance orders from a highly targeted subset of candidates. • Second, leveraging the constructed dataset, we train a novel cross-modal reranker, STeReO. It is fine-tuned using Low-Rank Adaptation (LoRA) and operates via a unified token-based scoring mechanism. Furthermore, the rich annotations within our dataset enable the proposed algorithm to be optimized across various learning objectives, including pointwise, pairwise, and listwise reranking formulations. • Finally, we comprehensively evaluate STeReO on a cross-modal benchmark constructed from the Spoken SQuAD and MS MARCO datasets. Extensive experiments across various backbone architectures demonstrated that our proposed method outperforms existing single-modality baselines. 2 The proposed method STeReO In this section, we present the proposed method, detailing the construction of a dataset equipped with ranking labels across heterogeneous modalities. Figure 2: Overview of the proposed framework STeReO. (a) construction dataset for mixed modality reranker, (b) training reranker, and (c) inference procedure based on the proposed reranker model. Framework. Prior to detailing the proposed method, we establish the core framework of this study. Let q denote a text query and mD^m represent heterogeneous databases, where m∈M=text,speechm∈ M=\text,speech\. For each modality m, a modality-specific retriever RmR^m extracts a candidate set m=c1m,…,ckmC^m=\c_1^m,…,c_k^m\, where k is the number of candidates from the retriever. Subsequently, a reranker ℛR processes the union of these sets to yield the mixed evidence ⋆=ℛ(⋃m∈Mm)C =R( _m∈ MC^m). Finally, the Audio Language Model (ALM) generates the answer a to the query as a=ALM(⋆,q)a=ALM(C ,q). The inference procedure is described in Figure 2 (c). Here, our primary focus is on training the reranker module ℛR. To achieve this, we first construct the necessary training dataset (Phase 1), followed by the formal training of ℛR (Phase 2). 2.1 Phase 1: Dataset Construction In Phase 1, we construct a cross-modal dataset by aggregating candidates from independent, modality-specific retrievers and annotating them with unified relevance labels using a foundation ALM, e.g., GPT (Figure 2 (a)). Domain-specific Retrieving. We first generate a candidate set mC^m for each modality and compute initial relevance scores. To ensure high-quality retrieval in each domain, we employ specialized pre-trained retrievers: e5-mistral-7b-instruct [28] for text and a HuBERT-based SpeechRAG [18] for speech. Normalization and Fusion. Since independent retrievers operate on disparate score scales, direct comparison is infeasible. To facilitate efficient candidate selection and minimize downstream labeling costs, we align these scores using Z-normalization: r~i=ri−μmσmform∈M, r_i= r_i- _m _m m∈ M, Here, rir_i and r~i r_i represent the raw and normalized retrieval scores for the ithi^th candidate, while μm _m and σm _m denote the mean and standard deviation of scores for modality m given a query q. We then merge the candidates based on these normalized scores to form a unified labeling set L, from which the top-k samples are selected. Cross-modal Labeling via Foundation ALM. To establish a gold-standard ranking across different modalities, we utilize a foundation ALM (e.g., gpt-4o-audio-preview) as a unified evaluator. The top-k candidates in L=c1,…,ckL=\c_1,…,c_k\ are fed into the ALM, which assigns individual relevance scores y=y1,…,yky=\y_1,…,y_k\ by considering both audio and text contexts simultaneously. Detailed prompts and implementation specifications are provided in Section 3 and the Supplementary Material. 2.2 Phase 2: Reranker Training In phase 2, we fine-tune the reranker using the cross-modal candidate-relevance tuples (q,ci,yi)(q,c_i,y_i) generated in the previous phase. To ensure parameter efficiency, we employ Low-Rank Adaptation (LoRA) [12]. The training process is depicted in Figure 2 (b). Autoregressive Relevance Scoring. We adopt decoder-based ALMs, such as Ultravox [11], Qwen-Audio-Chat [8], and Qwen2-Audio [7], as our base architecture for reranker. Following established autoregressive reranking paradigms [26, 21], the model is trained to generate a scalar relevance score sis_i, based on the logit difference between the tokens Yes and No at the final token position: si=logit(Yes)−logit(No).s_i=logit( Yes)-logit( No). Table 1: Training objectives for reranker optimization. σ(⋅)σ(·) denotes the sigmoid, and τ is a hyperparameter set to 1010. Method Objective Formulation Pointwise ℒpoint=−1|⋆|∑i[yilogσ(si)+(1−yi)log(1−σ(si))] _point=- 1|C | _i [y_i σ(s_i)+(1-y_i) (1-σ(s_i)) ] Pairwise ℒpair=1||∑(i,j)∈(q)log(1+exp(−δij⋅(si−sj)))where=(i,j):ci,cj∈⋆,yi≠yj,δij=sign(yi−yj) aligned L_pair=& 1|P| _(i,j) (q) (1+ (- _ij·(s_i-s_j) ) )\\ &where =\(i,j):c_i,c_j ,y_i≠ y_j\, _ij=sign(y_i-y_j) aligned Listwise ℒlist=1−1IDCG∑i2yi−1log2(1+r^i)wherer^i=1+∑j≠iσ(τ(sj−si)) aligned L_list=1- 1IDCG _i 2^y_i-1 _2(1+ r_i) r_i=1+ _j≠ iσ (τ(s_j-s_i) ) aligned Table 2: Reranking performance under Single- and Mixed-domain settings. All results use Max pooling for audio window size. Single-domain Mixed-domain Spoken SQuAD MS MARCO Spoken SQuAD MS MARCO Backbone Obj. Hit@1 MRR NDCG@5 Hit@1 MRR NDCG@5 Hit@1 MRR NDCG@5 Hit@1 MRR NDCG@5 Baseline Z-score 0.6927 0.7498 0.7678 0.6864 0.7694 0.7987 0.5273 0.6408 0.6857 0.6621 0.7510 0.7846 Ultravox Pointwise 0.7834 0.8008 0.8057 0.7432 0.8060 0.8262 0.7630 0.7892 0.7971 0.6977 0.7755 0.8032 Pairwise 0.4855 0.6288 0.6775 0.4394 0.6195 0.6865 0.3602 0.5434 0.6131 0.3336 0.5349 0.6220 Listwise 0.4430 0.6017 0.6572 0.3356 0.5487 0.6332 0.4430 0.6017 0.6572 0.1185 0.3737 0.4996 Qwen-Audio-Chat Pointwise 0.7644 0.7904 0.7980 0.7244 0.7951 0.8180 0.7334 0.7728 0.7849 0.6907 0.7720 0.8006 Pairwise 0.7561 0.7861 0.7948 0.7279 0.7971 0.8196 0.7174 0.7635 0.7780 0.7076 0.7834 0.8092 Listwise 0.7361 0.7755 0.7870 0.6967 0.7792 0.8062 0.7275 0.7703 0.7831 0.5898 0.7128 0.7564 Qwen2-Audio Pointwise 0.7753 0.7966 0.8026 0.7337 0.8004 0.8220 0.7584 0.7872 0.7957 0.6997 0.7777 0.8049 Pairwise 0.7705 0.7937 0.8005 0.7405 0.8047 0.8252 0.7467 0.7803 0.7905 0.7164 0.7885 0.8130 Listwise 0.7580 0.7870 0.7955 0.6872 0.7744 0.8027 0.7557 0.7858 0.7946 0.4047 0.5976 0.6701 Training Objectives. We optimize the reranker using the objective functions detailed in Table 1. Our framework is designed for high flexibility, supporting Pointwise (binary cross-entropy) [17, 19], Pairwise (RankNet-style) [3, 4], and Listwise (ApproxNDCG) [5, 29, 20, 25] loss functions. This modularity allows any of these widely adopted objectives to be seamlessly integrated into our training pipeline. Audio Windowing and Score Aggregation. To handle potential evidence localization within long audio passages, we segment each candidate into W fixed-duration windows. During training, a single window is randomly sampled for scoring to serve as a stochastic regularizer, thereby enhancing model robustness. In contrast, during inference, the windows are scored independently, with the final passage-level score sis_i obtained by aggregating these window-level scores via mean(⋅) mean(·) or max(⋅) max(·). 3 Experiment 3.1 Experimental Setup This section details the experimental setup and evaluation metrics used in our evaluation. Datasets. We evaluate STeReO using a fixed top-k pipeline (k=5k=5) on two distinct datasets: Spoken SQuAD [15] consisting of text queries and TTS-generated audio passages, and MS MARCO [2], comprising text-only web passages. The combined retrieval pool contains approximately 2.82.8K audio and 9.19.1K text passages. These corpora feature non-overlapping passages and minimal query overlap. Given that the datasets differ in both domain and modality, this setup provides a challenging heterogeneous environment to assess the robustness of cross-modal reranking beyond simple modality-based discrimination. Models. Our experiments evaluate three audio-native backbone architectures as student rerankers: Ultravox [11], Qwen-Audio-Chat [8], and Qwen2-Audio [7]. We compare these three against Z-score retrieval as the primary baseline, which ranks candidates through the modality-wise normalization of retriever scores as described in Section 2. Table 3: Evaluation of the foundation ALM. Label quality is measured on 55K samples, and ranking performance is assessed via Hit@1 on 11K, respectively. Category Metric Value Gain (Δ ) Label Quality F1 Score 0.700 - MCC 0.614 - Ranking (Hit@1) Z-score (Baseline) 0.578 - Foundation ALM (gpt-4o) 0.652 +0.074 Oracle (GT) 0.825 +0.247 Dataset Annotation. We use gpt-4o-audio-preview11 1 Since gpt-4o-audio-preview does not support text-only queries, gpt-4o is used for the candidates that consist solely of text. for dataset annotation, employing deterministic decoding and structured output to ensure annotation consistency. To ensure the reliability of the generated labels, we evaluate the gpt-4o-audio-preview model’s performance on a held-out set of 5,0005,000 candidates. As shown in Table 3, the foundation ALM achieves an F1 score of 0.700 and an MCC (Matthews Correlation Coefficient) of 0.614 against passage-ID ground truth, confirming high-quality label synthesis. Moreover, we assess the ranking signal by evaluating the foundation ALM’s direct ranking performance across 1,0001,000 queries. The annotated labels achieve a Hit@1 of 0.652, outperforming Z-score retrieval baseline by 0.0740.074. Training. To evaluate the architectural modularity of the proposed algorithm, we train the reranker under three distinct objectives: Pointwise, Pairwise, and Listwise as denoted in Table 1. All models are fine-tuned via LoRA with hyperparameter r=16r=16, α=32α=32, and dropout probability 0.050.05 for 33 epochs using the AdamW optimizer. We set the learning rate to 2×10−42× 10^-4 with 10−210^-2 weight decay, a 10%10\% linear warmup, and a gradient accumulation factor of 22. Following the audio windowing strategy, as described in Section 2, we utilize 30 second segments (i.e., W=4W=4) within a 120 second total budget. Evaluation Scenario. To evaluate the model’s precision and robustness, we report results in two scenarios based on domain constraints applied after reranking. The Single-domain setting is designed to simulate a single-modality environment, ensuring that our approach maintains high performance even with a specific domain by eliminating cross-domain noise. In contrast, the Mixed-domain setting evaluates the model against the full heterogeneous pool. This scenario assesses the model’s ability to discriminate relevance in complex, multi-modal contexts where candidates from diverse sources are presented simultaneously. Evaluation Metric. We assess reranking performance using Hit@1, MRR (Mean Reciprocal Rank), and NDCG@5 (Normalized Discounted Cumulative Gain), which measure the model’s ability to correctly identify the ground-truth passage ID within the top-k candidates. These metrics evaluate the accuracy of the ranked list by checking if the target passage is successfully retrieved at the top positions. For downstream QA tasks, we additionally report Exact Match (EM) based on substring matching to evaluate the fidelity of the generated answers against the reference text. Both Single-domain and Mixed-domain scenarios utilize an identical set of held-out evaluation queries (∼ 13K from SQuAD and 8.78.7K from MS MARCO). Table 4: Downstream QA performance (EM) comparison. GPT evaluation is conducted on a representative subset of 11K samples. Spoken SQuAD MS MARCO Generator Scoring Model Single Mixed Generator Scoring Model Single Mixed Ultravox Z-score Retrieval 0.4555 0.3565 Ultravox Z-score Retrieval 0.3637 0.3531 Qwen2-Audio 0.5046 0.4762 Qwen2-Audio 0.3834 0.3729 Qwen-Audio-Chat 0.4986 0.4659 Qwen-Audio-Chat 0.3795 0.3666 Ultravox 0.5088 0.4796 Ultravox 0.3880 0.3719 Qwen2-Audio Z-score Retrieval 0.3335 0.2646 Qwen2-Audio Z-score Retrieval 0.3885 0.3774 Qwen2-Audio 0.3669 0.3468 Qwen2-Audio 0.4090 0.3981 Qwen-Audio-Chat 0.3633 0.3414 Qwen-Audio-Chat 0.4028 0.3898 Ultravox 0.3688 0.3513 Ultravox 0.4150 0.3976 Qwen-Audio-Chat Z-score Retrieval 0.3426 0.2794 Qwen-Audio-Chat Z-score Retrieval 0.3637 0.3516 Qwen2-Audio 0.3709 0.3499 Qwen2-Audio 0.3798 0.3742 Qwen-Audio-Chat 0.3656 0.3455 Qwen-Audio-Chat 0.3761 0.3688 Ultravox 0.3733 0.3555 Ultravox 0.3873 0.3705 GPT-4o-Audio-Preview Z-score Retrieval 0.6273 0.4960 GPT-4o-Audio-Preview Z-score Retrieval 0.3680 0.3590 Qwen2-Audio 0.6731 0.6320 Qwen2-Audio 0.3680 0.3660 Qwen-Audio-Chat 0.6731 0.6230 Qwen-Audio-Chat 0.3630 0.3560 Ultravox 0.6794 0.6350 Ultravox 0.3820 0.3650 3.2 Results Main results. Table 2 reports the reranking performance, demonstrating that our approach effectively maintains high precision in a single-domain environment (Single) while exhibiting robust discrimination in heterogeneous contexts (Mixed). In the Single setting, all three backbones successfully identify target information within their native domains, consistently outperforming the Z-score baseline when optimized with the appropriate objective. Notably, this performance advantage extends to the Mixed setting, where the models must distinguish relevance across a complex pool of combined speech and text candidates. While the pointwise objective yields the most stable results across backbones, the listwise approach shows higher sensitivity, with Ultravox experiencing a significant performance drop (Δ of 0.257) on MS MARCO compared to its pointwise counterpart. Downstream QA Performance. As illustrated in Table 4, we verify that the improvements in reranking precision directly translate into enhanced end-task performance for downstream Question Answering (QA). By feeding the top-11 passage from each pointwise STeReO into three ALM generators, we observe that our reranking framework consistently boosts substring-match EM over the Z-score baseline across all generator-dataset pairs. This enhancement is robustly maintained in both single-domain and mixed-domain scenarios, confirming that the reranker provides high-quality, relevant context that reduces potential hallucinations in generators. Notably, the most substantial gain is recorded in the Spoken SQuAD mixed setting, where the Ultravox generator achieves a 0.1230.123 EM accuracy increase rising from 0.3570.357 to 0.4800.480. These results demonstrate the practical utility of our mixed-modality based reranking approach in supporting accurate answer generation within complex, heterogeneous retrieval environments. 3.3 Analysis Table 5: The ratio of speech evidence in top-5 candidates. Query Set w/o Z-score (%) w/ Z-score (%) Spoken SQuAD 4.0 45.2 MS MARCO 0.1 18.3 Impact of Z-score Aggregation. As described in Table 5, to isolate retrieval-stage fusion effects, we compare unnormalized fusion, modality-wise Z-score, and Reciprocal Rank Fusion (RRF). On Spoken SQuAD, both unnormalized fusion and RRF suffer from modality collapse, yielding near-zero Hit@1 (≈0.05≈ 0.05) as raw text scores overwhelm speech candidates. Z-score normalization effectively resolves this bias, significantly recovering Hit@1 to 0.5270.527. While unnormalized fusion slightly outperforms Z-score on MS MARCO (0.6870.687 vs. 0.6620.662) due to the inherent dominance of text candidates, we adopt Z-score as the default. Unlike other methods, Z-score consistently ensures speech visibility across both domains, preventing the exclusion of speech candidates in mixed modality retrieval. Table 6: MRR under inference-time Mean/Max pooling. Backbone MS MARCO Spoken SQuAD Mean Max Mean Max Ultravox 0.7918 0.7755 0.7460 0.7892 Qwen-Audio-Chat 0.7821 0.7720 0.7286 0.7728 Qwen2-Audio 0.7885 0.7777 0.7468 0.7872 Analysis of Pooling Strategy. Table 6 compares Mean and Max pooling strategies for aggregating speech window scores in the pointwise Mixed setting. The choice of pooling primarily affects audio-heavy contexts; on MS MARCO, where the audio candidate share is relatively low, the performance gap is marginal, with Mean pooling holding a slight edge. However, on Spoken SQuAD where speech candidates constitute a larger portion of the top-55 pool, Max pooling consistently outperforms Mean across all backbone cases. Table 7: Window-length analysis on Qwen2-Audio (Mixed). Spoken SQuAD MS MARCO Window Hit@1 MRR NDCG@5 Hit@1 MRR NDCG@5 120s×1 0.7509 0.7822 0.7919 0.7092 0.7834 0.8091 60s×2 0.7515 0.7827 0.7922 0.7138 0.7862 0.8112 30s×4 0.7584 0.7872 0.7957 0.6997 0.7777 0.8049 Window-length Analysis. Table 7 compares three window configurations on Qwen2-Audio under a fixed 120 seconds audio budget and Max pooling. Given the inherent maximum sequence length constraints of the model’s audio encoder, segmenting long audio into multiple windows is necessary to capture full temporal information without truncation. On Spoken SQuAD, finer segmentation (30s×4) achieves the highest Hit@1 score (0.758), outperforming the single-window (120s×1). While the performance gap remains marginal on MS MARCO due to the lower audio candidate density, these results indicate that finer-grained windows can more effectively pinpoint salient information while operating within the model’s architectural limits. 4 Conclusion We investigated reranking in heterogeneous pools of speech and text, demonstrating that STeReO effectively maintains high precision in a single-domain environment while ensuring robust discrimination in mixed-modality scenarios. By addressing structural challenges through modality-wise Z-score normalization and an optimized audio windowing strategy to overcome sequence length constraints, we achieve consistent performance gains. Our findings highlight that the pointwise objective is the primary driver of stability, whereas listwise approaches can lead to sharp degradation in heterogeneous settings. Ultimately, these reranking improvements translate into enhanced downstream QA performance, providing reliable context for accurate answer generation. While this study utilizes TTS-based data, i.e., Spoken SQuAD, future work will focus on validating these insights with natural, spontaneous speech. 5 Acknowledgments This work was supported by the Institute of Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (RS-2025-25464461, AI’s Vision of Harmony: A Fair and Transparent Multimodal Agentic Platform for Conflict Mediation) 6 Generative AI Use Disclosure Large Language Model assistance was used for language editing and polishing of portions of this manuscript. Beyond the specific experimental procedures explicitly described in the methodology, such as the use of generative models for synthetic dataset construction, no generative AI tool was used to produce experimental results, figures, tables or the scientific content of this work. References [1] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. (2023) Gpt-4 technical report. arXiv preprint arXiv:2303.08774. Cited by: §1. [2] P. Bajaj, D. Campos, N. Craswell, L. Deng, J. Gao, X. Liu, R. Majumder, A. McNamara, B. Mitra, T. Nguyen, M. Rosenberg, X. Song, A. Stoica, S. Tiwary, and T. Wang (2018) MS marco: a human generated machine reading comprehension dataset. External Links: 1611.09268, Link Cited by: §3.1. [3] C. Burges, T. Shaked, E. Renshaw, A. Lazier, M. Deeds, N. Hamilton, and G. Hullender (2005) Learning to rank using gradient descent. In Proceedings of the 22nd International Conference on Machine Learning, ICML ’05, New York, NY, USA, p. 89–96. External Links: ISBN 1595931805, Link, Document Cited by: §2.2. [4] C. J. Burges (2010) From ranknet to lambdarank to lambdamart: an overview. Learning 11 (23-581), p. 81. Cited by: §2.2. [5] Z. Cao, T. Qin, T. Liu, M. Tsai, and H. Li (2007) Learning to rank: from pairwise approach to listwise approach. In Proceedings of the 24th international conference on Machine learning, p. 129–136. Cited by: §2.2. [6] Y. Chen, S. Ji, H. Wang, Z. Wang, S. Chen, J. He, J. Xu, and Z. Zhao (2025) WavRAG: audio-integrated retrieval augmented generation for spoken dialogue models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, p. 12505–12523. External Links: Document, Link Cited by: §1. [7] Y. Chu, J. Xu, Q. Yang, H. Wei, X. Wei, Z. Guo, Y. Leng, Y. Lv, J. He, J. Lin, et al. (2024) Qwen2-audio technical report. arXiv preprint arXiv:2407.10759. Cited by: §1, §2.2, §3.1. [8] Y. Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou (2023) Qwen-audio: advancing universal audio understanding via unified large-scale audio-language models. External Links: 2311.07919, Link Cited by: §2.2, §3.1. [9] B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang (2023) Clap learning audio concepts from natural language supervision. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1–5. Cited by: §1. [10] P. Feng, Z. Ma, W. Chen, Y. Li, S. Wang, K. Yu, and X. Chen (2025) Enhancing speech-to-speech dialogue modeling with end-to-end retrieval-augmented generation. arXiv preprint arXiv:2505.00028. Cited by: §1. [11] Fixie AI (2024) Ultravox: a fast multimodal llm for real-time voice. Note: https://github.com/fixie-ai/ultravoxAccessed: 2025 Cited by: §2.2, §3.1. [12] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2022) LoRA: low-rank adaptation of large language models. In International Conference on Learning Representations (ICLR), Cited by: §2.2. [13] L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, et al. (2025) A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems 43 (2), p. 1–55. Cited by: §1. [14] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih, T. Rocktäschel, et al. (2020) Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems 33, p. 9459–9474. Cited by: §1. [15] C. Li, S. Wu, C. Liu, and H. Lee (2018) Spoken squad: a study of mitigating the impact of speech recognition errors on listening comprehension. arXiv preprint arXiv:1804.00320. Cited by: §3.1. [16] V. W. Liang, Y. Zhang, Y. Kwon, S. Yeung, and J. Y. Zou (2022) Mind the gap: understanding the modality gap in multi-modal contrastive representation learning. Advances in Neural Information Processing Systems 35, p. 17612–17625. Cited by: §1. [17] T. Liu (2009) Learning to rank for information retrieval. Foundations and Trends® in Information Retrieval 3 (3), p. 225–331. Cited by: §2.2. [18] D. J. Min, K. Mundnich, A. Lapastora, E. Soltanmohammadi, S. Ronanki, and K. Han (2024) Speech retrieval-augmented generation without automatic speech recognition. External Links: 2412.16500, Link Cited by: §1, §2.1. [19] R. Nogueira, Z. Jiang, R. Pradeep, and J. Lin (2020) Document ranking with a pretrained sequence-to-sequence model. In Findings of the association for computational linguistics: EMNLP 2020, p. 708–718. Cited by: §2.2. [20] T. Qin, T. Liu, and H. Li (2010) A general approximation framework for direct optimization of information retrieval measures. Information retrieval 13 (4), p. 375–397. Cited by: §2.2. [21] Z. Qin, R. Jagerman, K. Hui, H. Zhuang, J. Wu, L. Yan, J. Shen, T. Liu, J. Liu, D. Metzler, et al. (2024) Large language models are effective text rankers with pairwise ranking prompting. In Findings of the Association for Computational Linguistics: NAACL 2024, p. 1504–1518. Cited by: §2.2. [22] Z. Rackauckas and J. Hirschberg (2025) VoxRAG: a step toward transcription-free RAG systems in spoken question answering. In Proceedings of the 1st Workshop on Multimodal Augmented Generation via Multimodal Retrieval (MAGMaR 2025), R. Kriz and K. Murray (Eds.), Vienna, Austria, p. 40–46. External Links: Link, Document, ISBN 979-8-89176-280-0 Cited by: §1. [23] K. Saijo, J. Ebbers, F. G. Germain, S. Khurana, G. Wichern, and J. Le Roux (2025) Leveraging audio-only data for text-queried target sound extraction. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1–5. Cited by: §1. [24] F. Schlatt, M. Fröbe, H. Scells, S. Zhuang, B. Koopman, G. Zuccon, B. Stein, M. Potthast, and M. Hagen (2025) Set-encoder: permutation-invariant inter-passage attention for listwise passage re-ranking with cross-encoders. In European Conference on Information Retrieval, p. 1–19. Cited by: §1. [25] S. Sharifymoghaddam, R. Pradeep, A. Slavescu, R. Nguyen, A. Xu, Z. Chen, Y. Zhang, Y. Chen, J. Xian, and J. Lin (2025) Rankllm: a python package for reranking with llms. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, p. 3681–3690. Cited by: §2.2. [26] W. Sun, L. Yan, X. Ma, S. Wang, P. Ren, Z. Chen, D. Yin, and Z. Ren (2023) Is chatgpt good at search? investigating large language models as re-ranking agents. In Proceedings of the 2023 conference on empirical methods in natural language processing, p. 14918–14937. Cited by: §2.2. [27] G. Team, R. Anil, S. Borgeaud, J. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al. (2023) Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805. Cited by: §1. [28] L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, and F. Wei (2024) Improving text embeddings with large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 11897–11916. Cited by: §2.1. [29] F. Xia, T. Liu, J. Wang, W. Zhang, and H. Li (2008) Listwise approach to learning to rank: theory and algorithm. In Proceedings of the 25th international conference on Machine learning, p. 1192–1199. Cited by: §2.2. [30] S. Yoon, E. Choi, J. Kim, H. Yun, Y. Kim, and S. Hwang (2024) Listt5: listwise reranking with fusion-in-decoder improves zero-shot retrieval. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 2287–2308. Cited by: §1.