Paper deep dive
Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt
Yanfeng Shi, Pengfei Cai, Jun Liu, Qing Gu, Nan Jiang, Lirong Dai, Ian McLoughlin, Yan Song
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/18/2026, 1:29:45 AM
Summary
The paper introduces TimePro-RL, a framework designed to improve the fine-grained temporal perception of Large Audio-Language Models (LALMs). It utilizes 'Audio-Side Time Prompt' (ASTP) to inject timestamp embeddings into audio feature sequences and employs Reinforcement Learning (RL) with an adaptive temporal reward mechanism to optimize temporal alignment performance. Experiments show significant gains in audio grounding, sound event detection, and dense audio captioning.
Entities (5)
Relation Signals (3)
TimePro-RL → utilizes → Audio-Side Time Prompt
confidence 98% · TimePro-RL framework for fine-grained temporal perception... we propose Audio-Side Time Prompt
TimePro-RL → appliesto → LALMs
confidence 95% · enhance the fine-grained temporal perception of LALMs via two primary modules
TimePro-RL → improves → Temporal Perception
confidence 95% · TimePro-RL framework for fine-grained temporal perception.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Audio-Language Models (LALMs) enable general audio understanding and demonstrate remarkable performance across various audio tasks. However, these models still face challenges in temporal perception (e.g., inferring event onset and offset), leading to limited utility in fine-grained scenarios. To address this issue, we propose Audio-Side Time Prompt and leverage Reinforcement Learning (RL) to develop the TimePro-RL framework for fine-grained temporal perception. Specifically, we encode timestamps as embeddings and interleave them within the audio feature sequence as temporal coordinates to prompt the model. Furthermore, we introduce RL following Supervised Fine-Tuning (SFT) to directly optimize temporal alignment performance. Experiments demonstrate that TimePro-RL achieves significant performance gains across a range of audio temporal tasks, such as audio grounding, sound event detection, and dense audio captioning, validating its robust effectiveness.
Tags
Links
- Source: https://arxiv.org/abs/2604.13715v1
- Canonical: https://arxiv.org/abs/2604.13715v1
Trouble viewing inline? Open PDF directly →
Full Text
27,674 characters extracted from source content.
Expand or collapse full text
Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt Yanfeng Shi 1 , Pengfei Cai 1 , Jun Liu 1 , Qing Gu 1 , Nan Jiang 1 , Lirong Dai 1 , Ian McLoughlin 2 , Yan Song 1,∗ 1 National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China 2 ICT Cluster, Singapore Institute of Technology, Singapore yanfshi@mail.ustc.edu.cn, lrdai@ustc.edu.cn, ian.mcloughlin@singaporetech.edu.sg, songy@ustc.edu.cn Abstract Large Audio-Language Models (LALMs) enable general audio understanding and demonstrate remarkable performance across various audio tasks. However, these models still face chal- lenges in temporal perception (e.g., inferring event onset and offset), leading to limited utility in fine-grained scenarios. To address this issue, we propose Audio-Side Time Prompt and leverage Reinforcement Learning (RL) to develop the TimePro- RL framework for fine-grained temporal perception. Specifi- cally, we encode timestamps as embeddings and interleave them within the audio feature sequence as temporal coordinates to prompt the model. Furthermore, we introduce RL following Su- pervised Fine-Tuning (SFT) to directly optimize temporal align- ment performance. Experiments demonstrate that TimePro-RL achieves significant performance gains across a range of audio temporal tasks, such as audio grounding, sound event detection, and dense audio captioning, validating its robust effectiveness. Index Terms: large audio-language model, fine-grained tem- poral perception, reinforcement learning 1. Introduction Audio conveys a wealth of information, ranging from human speech to environmental events, and serves as a fundamental modality for perceiving the world[1, 2]. Large Audio-Language Models (LALMs) have significantly advanced general audio understanding by integrating the linguistic reasoning of Large Language Models (LLMs) with audio encoders [3, 4]. These models have demonstrated strong versatility across a wide range of applications, such as acoustic scene classification, audio cap- tioning and audio question answering [5, 6, 7]. Based on large- scale cross-modal pre-training, LALMs exhibit a remarkable ability to comprehend diverse acoustic content and interact with users through flexible natural language, providing a unified in- terface for audio-related tasks. Despite these advancements, research indicates that current LALMs still exhibit shortcomings in fine-grained temporal un- derstanding [8]. In many practical applications, models are re- quired not only to interpret acoustic content but also to capture the temporal structure or event boundaries within it[9, 10]. Al- though LALMs excel at semantic recognition, they often strug- gle to precisely infer the onset and offset timestamps of spe- cific sound events. This gap exposes a key limitation in fine- grained temporal perception, which is essential for temporally grounded tasks such as audio grounding[11] and sound event detection[12]. ** indicates the corresponding author. Prior work has strengthened temporal perception through the construction of time-annotated datasets [13], yielding sig- nificant performance improvements. Other approaches intro- duce specialized time tokens [14] to represent the concept of time, enhancing temporal alignment during generative infer- ence. While such advancements are promising, two aspects merit further attention: 1) The audio input in LALMs lacks the explicit modeling of physical temporal cues, limiting the precise alignment between semantic content and its actual tem- poral coordinates. 2) Supervised Fine-Tuning (SFT) primarily focuses on semantic correctness, lacking optimization signals to address time-boundary prediction deviations. Notably, Re- inforcement Learning (RL) has shown potential in Video Tem- poral Grounding (VTG) [15, 16], where reward signals can be directly designed around temporal alignment metrics, offering valuable insights for analogous tasks in the audio domain. Motivated by these observations, we propose the TimePro- RL framework to enhance the fine-grained temporal perception of LALMs via two primary modules: temporal information in- tegration and training objective design. Specifically, we intro- duce Audio-Side Time Prompt to inject temporal coordinates into audio features, then employ SFT to instruct the model in utilizing these coordinates. Subsequently, we further develop RL post-training with an advantage-driven adaptive temporal reward to incorporate temporal alignment quality into the op- timization objective. Through the synergy of these strategies, TimePro-RL achieves significant performance gains across a range of audio temporal tasks. 2. Method In this section, we elaborate on the TimePro-RL framework, the overall schematic of which is illustrated in Figure 1. We first present the infusion of temporal cues into LALM’s audio input (Section 2.1), and then describe the RL post-training paradigm tailored for audio temporal tasks, focusing on reward design (Section 2.2). 2.1. Audio-Side Time Prompt Most LALMs predominantly rely on position embeddings, such as RoPE [17], to capture sequential structure. While these models can be trained to extract temporal information from se- quence inputs [18], directly inferring absolute timestamps re- mains challenging. For VTG task, a previous study has over- laid frame indices onto video frames to provide perceptible time references, thereby enhancing temporal localization per- formance [19]. The core of this approach lies in providing “time arXiv:2604.13715v1 [cs.SD] 15 Apr 2026 Large Language Model Text Embedding Layer Audio Encoder Timestamp Embedding Layer A T A A T A A T 0.08s0.16s0.24s t (a) Audio-Side Time Prompt(b) SFT & RL Post-Training · · T Audio Feature Timestamp Embedding Text Embedding A T Audio Token Timestamp Token T Text Token Adaptive Temporal Reward Ground-Truth Time Segments Predicted Time Segments R 1 R 2 R G Group Relative Policy Optimization Subset Data Label LALM Extract Time Segments RL Stage · T LALM Label Label Text Tokens Tokenize · T Predict Text Tokens Full DataSFT Stage · When does “a railroad crossing rings ” happen? “a railroad crossing rings”: [[0.1, 4.4], [7.3, 10.0]] · T ·T ·T · · Group Predictions Group Size Figure 1: Overview of our TimePro-RL framework. Timestamp Embeddings are interleaved within audio features as time prompt, followed by SFT and RL post-training. prompt” at the input level, effectively reducing reasoning diffi- culty and mitigating hallucinations. Inspired by this, we propose Audio-Side Time Prompt (ASTP), in which timestamps are encoded into embedding vec- tors and interleaved within the audio feature sequence to con- stitute LALM’s audio input. Specifically, we first extend the tokenizer with a set of Timestamp Tokens (e.g., <0.04>), each corresponding to a specific time point in seconds. During pre- processing, the audio input is partitioned into an audio token sequence based on its duration and the output frame rate of the audio encoder. Since each position in this sequence maintains a fixed mapping to the timeline, Timestamp Tokens can be in- serted into the sequence according to their corresponding time points, serving as explicit temporal coordinates to prompt the model. The following shows an example of the processed input at the maximum time resolution: <s><audio><AUDIO><0.04><AUDIO><0.08><AUDIO> <0.12><AUDIO><0.16>·</audio>When does "a railroad crossing rings" happen?</s> where <s> and </s> indicate the start and end of the sequence, while <audio> and </audio> demarcate the audio segment. <AUDIO> serves as a placeholder token, which is then replaced by the corresponding audio frame feature. In this example, the audio encoder’s output frame rate is 25 Hz, and tokens such as <0.04> represent the inserted Timestamp Tokens. Following this, Timestamp Tokens in the sequence are mapped to vector representations via the Timestamp Embed- ding Layer. To ensure the stability of the embedding space, we adopt an initialization strategy based on semantic priors: each Timestamp Embedding is initialized as the mean of the subword embeddings derived from tokenizing its corresponding numeri- cal string. Taking <0.04> as an example: E ⟨0.04⟩ = 1 |T (0.04)| X u∈T (0.04) e u (1) where E ⟨0.04⟩ indicates the embedding vector of Timestamp Token <0.04>, e u represents the pre-trained embedding of the token with ID u in the vocabulary, andT (0.04) is the set of to- ken IDs obtained by tokenizing the numerical string “0.04”. Adding time prompt to LALM’s input aligns well with its autoregressive architecture. This enables the model to re- trieve temporal information from neighboring Timestamp Em- beddings when attending to event cues within audio frames. Since the input structure is altered, SFT is required to guide the model in correctly understanding and utilizing this prompt. The semantic initialization strategy facilitates this by transfer- ring pre-trained knowledge from the original language model. Moreover, all E ⟨t⟩ parameters are frozen during training to pre- vent semantic drift. 2.2. RL for audio temporal tasks While LALM’s temporal perception can be manifested in its ability to localize sound events, the standard SFT objective remains misaligned. Consider a ground-truth segment [5.0 s, 6.0 s], a reasonable prediction of [4.9 s, 5.9 s] would be heavily penalized by token-level cross-entropy loss, which risks over- fitting and impairs generalization capability [20]. RL offers a viable path to address this problem by optimizing for evaluation metrics, and has been successfully applied in generation tasks like summarization by directly utilizing ROUGE or CIDEr as rewards [21, 22]. However, the potential of RL for audio tem- poral tasks has yet to be extensively explored. For further enhancing fine-grained temporal perception of LALMs, we extend the training process following SFT by adopting Group Relative Policy Optimization (GRPO) [23], which computes relative advantages through group sampling. To align the training objective with temporal alignment perfor- mance, we utilize the Event-based F1 score (Eb-F1) [24], an established metric in sound event detection, as the main reward (r main ). However, we observe in our experiments that since Eb- F1 is a threshold-based discrete metric, the limited group size in GRPO may result in identical rewards among sampled pre- dictions, which leads to advantage degeneration and diminishes data efficiency. To address this, we incorporate a continuous auxiliary re- Table 1: Performance comparison of TimePro-RL and zero-shot / finetuned baseline models on audio temporal tasks. Bold indicates the best results in each category. ModelScale Audio GroundingSound Event Detection Dense Audio Captioning R@0.5R@0.7R@0.9mIoUEb-F1METEOREb-F1 Zero-shot Qwen2-Audio7B9.25.13.311.93.411.23.0 Qwen2.5-Omni7B25.417.410.627.713.710.510.4 Finetuned Audio-Flaming23B37.027.619.043.38.925.712.7 TimeAudio7B75.761.236.557.8−20.437.4 Qwen2-Audio7B74.857.934.669.649.832.235.0 Qwen2.5-Omni7B74.059.834.169.948.931.335.2 Kimi-Audio7B76.160.034.570.650.931.232.7 Post-Trained with TimePro-RL (Ours) Qwen2-Audio7B78.864.038.172.958.435.339.8 Qwen2.5-Omni7B80.166.339.874.457.633.940.7 ward (r aux ), such as the mean Intersection over Union (mIoU), to provide smoother optimization signals.We develop an advantage-driven adaptive temporal reward mechanism: R = ( r main ⊙ r aux , if Var(r main ) < ε r main ,otherwise (2) where R indicates the group reward vector for advantage com- putation, Var(·) represents the variance operator, and ε is the variance threshold. If r main lacks discriminability among predic- tions in a group, the algorithm adopts the element-wise product of r main and r aux as the fused reward. This strategy leverages the smoothness of r aux to recover the advantage signal while utilizing r main as a weight to regulate optimization intensity. By dynamically adjusting the reward calculation based on real-time data conditions during training, this mechanism improves data efficiency while maintaining high temporal alignment quality. 3. Experimental setup 3.1. Tasks and datasets We conduct experiments across three representative audio tem- poral tasks: Audio Grounding (AG) [11]. The audio grounding task is defined as localizing a specific sound event within an audio clip based on a descriptive natural language query. For this task, we utilize the FTAR dataset [8] and require the model to output the corresponding timestamps in the"query": [onset, offset] format. Performance is measured via Intersection over Union (IoU) metrics, where we report recall values at vari- ous thresholds including R@0.5, R@0.7, and R@0.9 alongside the mean IoU (mIoU). Sound Event Detection (SED) [12]. Sound event detec- tion involves identifying event categories from a predefined set along with their respective occurrence periods. We conduct this task on the DESED dataset [25], with model outputs format- ted as"event": [onset, offset]. The evaluation metric of this task leverages the Eb-F1 score [24], and we set the boundary tolerance to 0.2 s. Dense Audio Captioning (DAC) [8]. For dense audio cap- tioning, the model should generate descriptions for the sound events in the audio clip paired with their associated timestamps. This task is constructed using the FTAR dataset [8], and out- puts follow the format: onset-offset, description. To measure performance, we employ METEOR [26] to assess the linguistic quality of the generated captions and Eb-F1 to evaluate the accuracy of temporal localization. The sample sizes for the training and test sets of each task are detailed in Table 2. Table 2: Summary of data statistics across the three tasks. TaskTraining Size Test Size Audio Grounding61,862483 Sound Event Detection15,0411,153 Dense Audio Captioning92,443741 3.2. Implementation details To evaluate the effectiveness of the TimePro-RL framework, we utilize Qwen2-Audio [27] and Qwen2.5-Omni [28] as base models. Both of them integrate Whisper [29] as the audio en- coder with an output frame rate of 25 Hz. To support Audio- Side Time Prompt at the maximum time resolution, we expand the tokenizer with 750 Timestamp Tokens, covering from 0 s to 30 s with a stride of 0.04 s. We conduct parameter-efficient fine-tuning using LoRA [30] with r = 8 and α = 32. The over- all post-training pipeline consists of two stages: 1) SFT: The models are trained on the full dataset for 3 epochs with a learn- ing rate of 1× 10 −5 . 2) RL: Following SFT, we take GRPO for only a single epoch using a subset of 10,200 samples. The group size is set to 4, and the learning rate is 1 × 10 −6 . In the adaptive temporal reward mechanism, we employ Eb-F1 as r main across all three tasks. For r aux , we utilize mIoU for AG and SED, while METEOR for DAC. The variance threshold ε is specified as 1× 10 −6 . 4. Results In this section, we first evaluate TimePro-RL against baselines under zero-shot and finetuned settings. Then, we conduct ab- lation studies to analyze the contributions of components in TimePro-RL. Table 3: Ablation study of different components. ASTP denotes Audio-Side Time Prompt; “random init” refers to random initialization of Timestamp Embeddings; RL(Eb-F1) indicates using only Eb-F1 as the reward. MethodAudio GroundingSound Event Detection Dense Audio Captioning (Qwen2.5-Omni)R@0.5R@0.7R@0.9mIoUEb-F1METEOREb-F1 SFT Baseline74.059.834.169.948.931.335.2 w/ ASTP (random init)73.257.232.868.846.031.433.3 w/ ASTP77.661.735.871.750.132.637.0 w/ ASTP + RL (Eb-F1)77.863.138.972.756.931.638.1 w/ ASTP + RL80.166.339.874.457.633.940.7 4.1. Main results As shown in Table 1, we first evaluate Qwen2-Audio and Qwen2.5-Omni under zero-shot conditions, where their perfor- mance on high-precision metrics (e.g., R@0.9 and Eb-F1) is no- tably constrained, revealing the limitations of existing general- purpose LALMs in fine-grained temporal perception. Sub- sequently, we compare models post-trained via TimePro-RL framework with several prominent LALMs adapted by SFT on the same dataset, including Qwen2-Audio, Qwen2.5-Omni, Audio-Flamingo2 [31] and Kimi-Audio [32]. For TimeAudio, which is only trained on the FTAR dataset, we adopt the results reported in its original publication [8]. The results show that TimePro-RL consistently outperforms these baselines across multiple evaluation metrics, and crucially, it demonstrates a strong competitive edge in high-precision localization. For in- stance, in AG task, Qwen2.5-Omni improves from 34.1 during the SFT stage to 39.8 on R@0.9. Similarly, in DAC task, its Eb-F1 score rises from 35.2 to 40.7, while Qwen2-Audio also achieves a 4.8-point improvement. These gains validate the ef- fectiveness of TimePro-RL in enhancing the fine-grained tem- poral perception of LALMs. 4.2. Ablation study We conduct a series of ablation experiments based on Qwen2.5- Omni to systematically investigate the contributions of each component in TimePro-RL framework, with the results summa- rized in Table 3. Compared to the SFT baseline, using random initialization in Audio-Side Time Prompt leads to performance regressions across most metrics, such as a 2.9-point drop in SED Eb-F1, because randomly initialized Timestamp Embeddings introduce extraneous noise into the audio feature sequence. In contrast, the semantic initialization strategy proposed in Sec- tion 2.1 allows the model to correctly interpret and leverage Timestamp Embeddings as temporal coordinates, offering per- formance gains including a 1.7-point increase in AG R@0.9. This underscores the critical importance of semantic initializa- tion for Audio-Side Time Prompt. Furthermore, our results demonstrate that even a modest volume of RL training significantly boosts model performance, with AG R@0.9 increasing from 35.8 to 39.8 compared to the SFT (w/ ASTP) stage. In terms of the reward design, the adaptive temporal reward mechanism in Section 2.2 provides more balanced overall gains compared to the Eb-F1 only con- figuration. While the latter causes an optimization imbalance where the METEOR score for DAC task declines from the 32.6 achieved in the SFT (w/ ASTP) stage to 31.6, our adaptive ap- proach recovers this score to 33.9 and pushes the final Eb-F1 to 40.7, facilitating more comprehensive data utilization. Figure 2: Visualization of attention weights on Timestamp Em- beddings. For the audio grounding query “a train horn honk- ing”, the mel-spectrogram (top) is aligned with the chronolog- ically arranged attention weights assigned to each Timestamp Embedding (bottom). Figure 2 provides a visual analysis of the attention weights assigned to Timestamp Embeddings to further elucidate the in- ternal mechanism of our framework. we extract the attention weights of the generate tokens toward each Timestamp Em- bedding from the model’s final layer and plot them along the timeline. The attention map reveals that the model’s focus ex- hibits high-intensity activations that precisely align with the on- set and offset boundaries of the sound events marked in the mel- spectrogram. This sharp concentration of attention on the start and end coordinates suggests that the model effectively lever- ages Timestamp Embeddings to capture precise temporal cues for sound events. Such clear alignment provides intuitive evi- dence for the interpretability of Audio-Side Time Prompt, con- firming that it successfully guides the model in perceiving and utilizing fine-grained temporal information integrated in the au- dio feature sequence. 5. Conclusion In this paper, we propose TimePro-RL, a framework to enhance the fine-grained temporal perception of LALMs. TimePro-RL interleaves Timestamp Embeddings into audio feature sequence to provide temporal cues, and introduces RL post-training with an adaptive temporal reward designed for temporal alignment to further strengthen temporal capabilities. Experiments show that TimePro-RL achieves significant performance gains across a range of audio temporal tasks, validating the effectiveness of the synergy between Audio-Side Time Prompt and RL post- training. Future research will explore the application of our method in complex reasoning scenarios, such as Chain-of- Thought (CoT), where fine-grained temporal cues serve as crit- ical intermediate evidence. 6. References [1] J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontol- ogy and human-labeled dataset for audio events,” in IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2017, p. 776–780. [2] Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumb- ley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, p. 2880–2894, 2020. [3] S. Deshmukh, B. Elizalde, R. Singh, and H. Wang, “Pengi: An audio language model for audio tasks,” in Advances in Neural In- formation Processing Systems (NeurIPS), vol. 36. Curran Asso- ciates, Inc., 2023, p. 18 090–18 108. [4] C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “SALMONN: towards generic hearing abilities for large language models,” in The Twelfth International Conference on Learning Representations (ICLR). OpenReview.net, 2024. [5] Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. R. Glass, “Listen, think, and understand,” in The Twelfth International Conference on Learning Representations (ICLR). OpenReview.net, 2024. [6] Y. Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou, “Qwen-audio: Advancing universal audio understand- ing via unified large-scale audio-language models,” arXiv preprint arXiv:2311.07919, 2023. [7] P. K. Rubenstein, C. Asawaroengchai, D. D. Nguyen, A. Bapna, Z. Borsos, F. d. C. Quitry, P. Chen, D. E. Badawy, W. Han, E. Kharitonov et al., “Audiopalm: A large language model that can speak and listen,” arXiv preprint arXiv:2306.12925, 2023. [8] H. Wang, Y. Li, S. Ma, H. Liu, and X. Wang, “Listening between the frames: Bridging temporal gaps in large audio-language mod- els,” arXiv preprint arXiv:2511.11039, 2025. [9] A. Mesaros, T. Heittola, T. Virtanen, and M. D. Plumbley, “Sound event detection: A tutorial,” IEEE Signal Processing Magazine, vol. 38, no. 5, p. 67–83, 2021. [10] C.-H. H. Yang, S. Ghosh, Q. Wang, J. Kim, H. Hong, S. Ku- mar, G. Zhong, Z. Kong, S. Sakshi, V. Lokegaonkar et al., “Multi-domain audio question answering toward acoustic con- tent reasoning in the DCASE 2025 challenge,” arXiv preprint arXiv:2505.07365, 2025. [11] X. Xu, H. Dinkel, M. Wu, and K. Yu, “Text-to-audio grounding: Building correspondence between captions and sound events,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, p. 606–610. [12] D. Stowell, D. Giannoulis, E. Benetos, M. Lagrange, and M. D. Plumbley, “Detection and classification of acoustic scenes and events,” IEEE Transactions on Multimedia, vol. 17, no. 10, p. 1733–1746, 2015. [13] Z. Xie, X. Xu, Z. Wu, and M. Wu, “Audiotime:A temporally-aligned audio-text benchmark dataset,” in IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, p. 1–5. [14] H. Wang, Z. Xu, Y. Cheng, S. Diao, Y. Zhou, Y. Cao, Q. Wang, W. Ge, and L. Huang, “Grounded-videollm: Sharpening fine- grained temporal grounding in video large language models,” in Findings of the Association for Computational Linguistics (EMNLP). Association for Computational Linguistics, 2025, p. 959–975. [15] X. Li, Z. Yan, D. Meng, L. Dong, X. Zeng, Y. He, Y. Wang, Y. Qiao, Y. Wang, and L. Wang, “Videochat-r1: Enhancing spatio-temporal perception via reinforcement fine-tuning,” arXiv preprint arXiv:2504.06958, 2025. [16] L. Dong, H. Zhang, H. Lin, Z. Yan, X. Zeng, H. Zhang, Y. Huang, Y. Wang, Z.-H. Ling, L. Wang et al., “Videotg-r1: Boosting video temporal grounding via curriculum reinforcement learning on re- flected boundary annotations,” arXiv preprint arXiv:2510.23397, 2025. [17] J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,” Neuro- computing, vol. 568, p. 127063, 2024. [18] A. K. Sridhar, Y. Guo, and E. Visser, “Enhancing temporal un- derstanding in audio question answering for large audio language models,” in Conference of the Nations of the Americas Chapter of the ACL (NAACL), 2025, p. 1026–1035. [19] Y. Wu, X. Hu, Y. Sun, Y. Zhou, W. Zhu, F. Rao, B. Schiele, and X. Yang, “Number it: Temporal grounding videos like flipping manga,” in IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), 2025, p. 13 754–13 765. [20] Y. Wang, Z. Wang, B. Xu, Y. Du, K. Lin, Z. Xiao, Z. Yue, J. Ju, L. Zhang, D. Yang et al., “Time-r1: Post-training large vision language model for temporal video grounding,” arXiv preprint arXiv:2503.13377, 2025. [21] R. Paulus, C. Xiong, and R. Socher, “A deep reinforced model for abstractive summarization,” in International Conference on Learning Representations (ICLR). OpenReview.net, 2018. [22] S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self- critical sequence training for image captioning,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017, p. 7008–7024. [23] D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi et al., “Deepseek-r1: Incentivizing reason- ing capability in llms via reinforcement learning,” arXiv preprint arXiv:2501.12948, 2025. [24] A. Mesaros, T. Heittola, and T. Virtanen, “Metrics for polyphonic sound event detection,” Applied Sciences, vol. 6, no. 6, p. 162, 2016. [25] N. Turpault, R. Serizel, A. P. Shah, and J. Salamon, “Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,” in Workshop on Detection and Classifica- tion of Acoustic Scenes and Events (DCASE), 2019. [26] S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the acl workshop on intrinsic and extrinsic eval- uation measures for machine translation and/or summarization, 2005, p. 65–72. [27] Y. Chu, J. Xu, Q. Yang, H. Wei, X. Wei, Z. Guo, Y. Leng, Y. Lv, J. He, J. Lin et al., “Qwen2-audio technical report,” arXiv preprint arXiv:2407.10759, 2024. [28] J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y. Fan, K. Dang et al., “Qwen2.5-omni technical report,” arXiv preprint arXiv:2503.20215, 2025. [29] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International Conference on Machine Learning (ICML). PMLR, 2023, p. 28 492–28 518. [30] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al., “Lora: Low-rank adaptation of large language models.” in International Conference on Learning Rep- resentations (ICLR), vol. 1, no. 2, 2022, p. 3. [31] S. Ghosh, Z. Kong, S. Kumar, S. Sakshi, J. Kim, W. Ping, R. Valle, D. Manocha, and B. Catanzaro, “Audio flamingo 2: An audio- language model with long-audio understanding and expert rea- soning abilities,” in International Conference on Machine Learn- ing (ICML), ser. Proceedings of Machine Learning Research, vol. 267. PMLR / OpenReview.net, 2025. [32] D. Ding, Z. Ju, Y. Leng, S. Liu, T. Liu, Z. Shang, K. Shen, W. Song, X. Tan, H. Tang et al., “Kimi-audio technical report,” arXiv preprint arXiv:2504.18425, 2025.