Paper deep dive
Nudging Hidden States: Training-Free Model Steering for Chain-of-Thought Reasoning in Large Audio-Language Models
Lok-Lam Ieong, Chia-Chien Chen, Chih-Kai Yang, Yu-Han Huang, An-Yu Cheng, Hung-yi Lee
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 3/22/2026, 5:12:11 AM
Summary
The paper introduces a training-free model steering framework to enhance Chain-of-Thought (CoT) reasoning in Large Audio-Language Models (LALMs). By manipulating hidden states during inference, the authors propose three strategies: Vanilla Steering (instance-specific), Speech-derived Generalized Steering (SGS), and Text-derived Generalized Steering (TGS). Results demonstrate that these methods, particularly TGS, achieve significant accuracy gains across multiple benchmarks and models, highlighting the effectiveness of cross-modal transfer and data efficiency in steering LALMs.
Entities (10)
Relation Signals (3)
Model Steering → enhances → Chain-of-Thought
confidence 95% · This work investigates model steering as a representation-level intervention to improve CoT reasoning in LALMs.
Vanilla Steering → requires → Hidden states
confidence 90% · In Vanilla Steering, the steering vector is constructed dynamically for each test sample at inference time.
TGS → transfersto → Speech-based reasoning
confidence 90% · TGS achieves higher average accuracy than CoT across all models despite deriving steering vectors purely from text data.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Chain-of-thought (CoT) prompting has been extended to large audio-language models (LALMs) to elicit reasoning, yet enhancing its effectiveness without training remains challenging. We study inference-time model steering as a training-free approach to improve LALM reasoning. We introduce three strategies using diverse information sources and evaluate them across four LALMs and four benchmarks. Results show general accuracy gains up to 4.4% over CoT prompting. Notably, we identify a cross-modal transfer where steering vectors derived from few text samples effectively guide speech-based reasoning, demonstrating high data efficiency. We also examine hyperparameter sensitivity to understand the robustness of these approaches. Our findings position model steering as a practical direction for strengthening LALM reasoning.
Tags
Links
- Source: https://arxiv.org/abs/2603.14636v1
- Canonical: https://arxiv.org/abs/2603.14636v1
Trouble viewing inline? Open PDF directly →
Full Text
33,775 characters extracted from source content.
Expand or collapse full text
Nudging Hidden States: Training-Free Model Steering for Chain-of-Thought Reasoning in Large Audio-Language Models Lok-Lam Ieong ∗ , Chia-Chien Chen ∗ , Chih-Kai Yang ∗ , Yu-Han Huang † , An-Yu Cheng † , Hung-yi Lee National Taiwan University, Taiwan chihkaiyang1124@gmail.com, hungyilee@ntu.edu.tw Abstract Chain-of-thought (CoT) prompting has been extended to large audio-language models (LALMs) to elicit reasoning, yet en- hancing its effectiveness without training remains challenging. We study inference-time model steering as a training-free ap- proach to improve LALM reasoning. We introduce three strate- gies using diverse information sources and evaluate them across four LALMs and four benchmarks. Results show general ac- curacy gains up to 4.4% over CoT prompting. Notably, we identify a cross-modal transfer where steering vectors derived from few text samples effectively guide speech-based reason- ing, demonstrating high data efficiency. We also examine hy- perparameter sensitivity to understand the robustness of these approaches. Our findings position model steering as a practical direction for strengthening LALM reasoning. Index Terms: large audio-language model, model steering, chain-of-thought, reasoning 1. Introduction Large audio-language models (LALMs) [1–8], which extend large language models (LLMs) [9–11] with auditory under- standing [12, 13], have recently achieved substantial progress. They demonstrate strong auditory perceptual capabilities [12, 14] and are increasingly regarded as a promising foundation for universal and interactive auditory intelligence [15–17]. How- ever, reasoning remains a fundamental limitation [18–20] that prevents these models from fully realizing this vision. Reasoning has long been central to artificial intelligence. In LLMs, Chain-of-Thought (CoT) prompting [21, 22] is a rep- resentative approach for eliciting structured, step-by-step rea- soning. Motivated by its success, recent work extends CoT to LALMs [23], with further improvements via supervised rea- soning data or reinforcement learning [24, 25]. However, these methods require additional supervision and substantial training cost. This raises a key question: Can we enhance CoT reason- ing in LALMs at inference time without extra training? Model steering [26–32] offers a training-free alternative by manipulating hidden states. In LLMs, steering vectors have been used for style control [26, 27], safety alignment [29], and performance improvement [31, 32]. In LALMs, steering has been applied to attribute recognition [12], hallucination mitiga- tion [33], and safety alignment [34], but its potential for enhanc- ing reasoning remains underexplored. This work investigates model steering as a representation- level intervention to improve CoT reasoning in LALMs. Reasoning-oriented steering directions are derived from the dif- ference between CoT and non-CoT hidden states and injected *† Equal Contribution. Figure 1: Example of reasoning enhanced by steering. during decoding. We propose three variants: Vanilla Steering, which extracts instance-specific vectors; Speech-derived Gen- eralized Steering (SGS), which extracts a shared vector from auxiliary spoken data; and Text-derived Generalized Steer- ing (TGS), which derives steering directions from text-only data and transfers them to speech-based reasoning. Figure 1 shows a qualitative example where steering improves intermediate rea- soning and corrects the final prediction. Extensive experiments on four advanced LALMs and four speech-based benchmarks show that steering generally im- proves CoT performance, with up to 4.4% absolute accuracy gains over CoT. Vanilla steering surpasses self-consistency [35] under a comparable computational budget while requiring fewer decoding operations. SGS and TGS demonstrate that effective steering vectors can be extracted without instance-specific ac- cess, achieving competitive improvements across models. No- tably, TGS achieves higher average accuracy than CoT across all models despite deriving steering vectors purely from text data. Hyperparameter sensitivity and data-efficiency analyses further show that generalized steering methods are more stable than instance-specific steering, with TGS requiring fewer sam- ples to reach competitive performance, highlighting its stability and data efficiency for enhancing speech-based reasoning. In summary, this work (1) introduces a training-free steer- ing framework for enhancing CoT reasoning in LALMs, (2) demonstrates its effectiveness and computational efficiency across multiple models and benchmarks, (3) reveals the feasibil- ity of generalized and cross-modal steering directions, and (4) provides empirical insights into the stability and data-efficiency characteristics of steering-based interventions. 2. Methodology Model steering consists of two phases: (1) an extraction phase, where steering vectors are derived, and (2) an injection phase, where the vectors are applied during generation. We describe arXiv:2603.14636v1 [cs.SD] 15 Mar 2026 Figure 2: Overview of our three proposed methods in the extraction phase, along with the subsequent injection phase. these phases and present three extraction methods (Figure 2). 2.1. Extraction Phase Given a model with L layers, the extraction phase constructs steering vectors from hidden states h (ℓ) t (s) ∈R d at layer ℓ ∈ [1,L] and token position t for input s. Following prior work, we extract steering directions from the last k layers [36–39] at the final prompt token [29, 40–42], denoted by ̄ h (ℓ) (s) for ℓ ∈ [L−k + 1,L]. Below, we introduce three extraction strategies. 2.1.1. Vanilla Steering In Vanilla Steering, the steering vector is constructed dynami- cally for each test sample at inference time. For a given sample, let a denote the audio input, x the task instruction, and p a fixed short chain-of-thought cue. We form two inputs: s cot = [a; x; p], s norm = [a; x]. (1) For each selected layer ℓ, the steering vector is defined as v (ℓ) vanilla = ̄ h (ℓ) (s cot )− ̄ h (ℓ) (s norm ).(2) Since v (ℓ) vanilla depends only on the hidden states under differ- ent prompting conditions for the same input, no ground truth or external supervision are involved. The procedure therefore constitutes a valid training-free inference-time intervention. 2.1.2. Speech-derived Generalized Steering (SGS) A limitation of Vanilla Steering is its computational overhead: extracting a sample-specific steering vector requires additional forward passes for each test input. This motivates constructing a shared steering direction that can be reused across samples. To this end, we propose Speech-derived Generalized Steer- ing (SGS). An external auxiliary spoken datasetD s ext is used to compute a shared steering vector applied uniformly to all test samples. For each example i ∈ D s ext , let a (i) denote the audio input and x (i) the task instruction. We construct s (i) cot = [a (i) ; x (i) ; p], s (i) norm = [a (i) ; x (i) ].(3) Following the Difference-in-Means paradigm [43–45], we aver- age the differences overD s ext to obtain a shared steering vector: v (ℓ) SGS = 1 |D s ext | X i∈D s ext ̄ h (ℓ) (s (i) cot )− ̄ h (ℓ) (s (i) norm ) .(4) Unlike Vanilla Steering, v (ℓ) SGS is computed once and reused across test sets. We evaluate whether such a shared direction can serve as a general reasoning-oriented steering signal. 2.1.3. Text-derived Generalized Steering (TGS) While SGS relies on spoken data for extraction, such data may be less accessible than text in practice. We therefore intro- duce Text-derived Generalized Steering (TGS), which derives a shared steering direction from text-only data and examines whether it can transfer to spoken reasoning tasks. Using an external textual datasetD t ext , for each example i∈ D t ext , let t (i) and x (i) denote the textual input and instruction, respectively. We construct s (i) cot = [t (i) ; x (i) ; p], s (i) norm = [t (i) ; x (i) ].(5) The shared steering vector is then computed in the same manner as SGS, by applying the Difference-in-Means overD t ext : v (ℓ) TGS = 1 |D t ext | X i∈D t ext ̄ h (ℓ) (s (i) cot )− ̄ h (ℓ) (s (i) norm ) .(6) The resulting vector is extracted entirely from text-only inputs and then transferred to spoken reasoning tasks at inference time. This setup allows us to assess whether a text-derived steering di- rection can improve CoT performance in speech-based settings. 2.2. Injection Phase After obtaining the steering vector v (ℓ) from the extraction phase, we scale it by a coefficient α, which controls the steering strength. The scaled vector is injected at inference time into the same set of layers selected during extraction, and the interven- tion is applied throughout decoding for all token positions. Let h (ℓ) t ∈R d denote the original hidden state at layer ℓ and token position t during inference. For each selected layer, we modify the hidden state by ̃ h (ℓ) t = h (ℓ) t + α v (ℓ) .(7) Following common practice in steering-based interventions, we apply norm-preserving injection to improve stability by rescal- ing the modified hidden state to match the original ℓ 2 norm: ˆ h (ℓ) t = ̃ h (ℓ) t · ∥h (ℓ) t ∥ 2 ∥ ̃ h (ℓ) t ∥ 2 .(8) Table 1: Accuracies (%) on four evaluation benchmarks and their micro-average (ALL). Bold values denote the best result among the five settings for each model on each benchmark. “Hyperparams.” specifies the last k layers and the scaling factor α used for each configuration. ∆ indicates the difference in average accuracy between our steering methods and the CoT baseline. ModelMethodHyperparams. (k, α)College (↑) High School (↑) Elementary (↑) ReveAL-CoT (↑)ALL (↑) ∆ (↑) Voxtral Normal–29.036.756.953.547.5– CoT–26.342.067.443.050.7– Vanilla (Ours)5, 0.02532.043.368.856.055.0+4.3 SGS (Ours)3, 0.129.343.965.857.053.8+3.1 TGS (Ours)4, 0.02526.341.367.154.552.8+2.1 Phi4-m Normal–26.024.833.138.531.1– CoT–36.041.969.668.557.9– Vanilla (Ours)5, 0.02535.047.069.169.559.3+1.4 SGS (Ours)3, 0.0530.047.068.870.558.9+1.0 TGS (Ours)3, 0.02530.048.971.470.560.4+2.5 Qwen2.5 Normal–40.045.652.151.548.8– CoT–65.079.983.071.077.7– Vanilla (Ours)3, 0.02556.079.683.974.077.6-0.1 SGS (Ours)4, 0.02565.080.784.771.578.7+1.0 TGS (Ours)3, 0.0564.081.983.671.578.5+0.8 AF3 Normal–26.128.226.046.531.0– CoT–27.424.230.445.531.5– Vanilla (Ours)3, 0.120.022.237.647.533.4+1.9 SGS (Ours)1, 0.02524.021.136.042.531.9+0.4 TGS (Ours) 5, 0.120.024.142.946.535.9+4.4 During inference, ˆ h (ℓ) t replaces h (ℓ) t at all token positions for the selected layers, and the model proceeds with standard for- ward computation and decoding using the modified states. LALMs may not always reliably follow CoT instructions, as their instruction-following ability can be weaker after multi- modal training [46]. Consequently, CoT prompts may not fully induce structured reasoning during decoding. By injecting a steering vector derived from CoT-induced state differences, we reinforce CoT-related activations in the hidden states, encour- aging the model to produce more structured reasoning. 3. Experimental Setups 3.1. Models We conducted our study in four advanced LALMs: Voxtral- mini-3B (Voxtral) [1],Phi4-Multimodal-Instruct (Phi4- m) [2], Qwen2.5-Omni-7B (Qwen2.5) [3], and Audio Flamingo 3 (AF3) [4].Unless otherwise specified, greedy decoding is used for all models. 3.2. Baselines We compare our steering methods with several baselines. Nor- mal denotes the default performance of the LALMs, while CoT [21, 22] refers to direct chain-of-thought prompting. We additionally include self-consistency [35], where each model generates three outputs with temperature 0.5 and ag- gregates them via majority voting. This baseline approximates the computational cost of Vanilla Steering, which also requires three forward passes per instance. However, self-consistency remains more expensive because all passes involve full genera- tion, whereas the extraction phase in Vanilla Steering does not. 3.3. Datasets and Evaluation Benchmarks We use BeyondAIME [47] (100 samples) as the external dataset for SGS and TGS, strictly disjoint from all evaluation bench- Table 2: Self-consistency (Self-Con) vs. vanilla steering on Col- lege (Col), High School (HS), Elementary (Elem) math, and ReveAL-CoT (RC), with overall accuracy (ALL). ModelMethodColHSElemRCALL Voxtral Self-Con25.036.366.951.050.4 Vanilla32.043.368.856.055.0 Phi4-m Self-Con 43.045.265.665.057.3 Vanilla35.047.069.169.559.3 Qwen2.5 Self-Con58.074.181.866.573.8 Vanilla56.079.683.974.077.6 AF3 Self-Con 27.023.739.447.535.3 Vanilla20.022.237.647.533.4 marks. For SGS, math questions are verbalized and synthesized with IndexTTS2 [48] to construct D s ext (with manual quality checks). For TGS, the original BeyondAIME formsD t ext . We tune the number of steered last layers k and the scal- ing factor α on the spoken GSM8K benchmark [49] from SpeechR [50], used as a development set disjoint from all eval- uation benchmarks. We search over α ∈ [0.025, 0.2] and k ∈ 1,..., 5, and report test performance using the configu- ration that achieves the best development accuracy. We evaluate steering on four spoken reasoning benchmarks: College, High School, and Elementary Mathematics from Vox- Eval [51], covering math problems of varying difficulty, and ReveAL-CoT from SpeechR [52], targeting spoken scientific reasoning. We follow their official evaluation protocols. 4. Results 4.1. Can Model Steering Improve Chain-of-thought? Table 1 presents the results. Steering improves over CoT in most settings, with 11 of 12 model–method combinations show- (a) Effect of the scaling factor α. (b) Effect of the number of steered last k layers. Figure 3: Hyperparameter sensitivity of the steering methods. ing positive gains in average accuracy. AF3 and Voxtral achieve the largest improvements (+4.4% and +4.3%), while other mod- els also benefit from multiple steering variants. Overall, these results suggest that representation interventions can systemati- cally alter CoT outcomes, with effects varying across models. We compare Vanilla Steering with self-consistency under a comparable computational setup. While both methods use three forward passes, self-consistency performs three full generation processes per instance, whereas Vanilla Steering requires only a single generation pass after extraction. As shown in Table 2, Vanilla Steering achieves higher overall accuracy on three of the four models. These results indicate that, under a matched forward-pass budget, steering typically attains better accuracy while requiring fewer generation processes. Beyond Vanilla Steering, we examine two variants, SGS and TGS, which extract general steering vectors without instance-specific information. Both methods improve average accuracy on all models and, in several cases, match or even out- perform Vanilla Steering (e.g., TGS on Phi-4-m and AF3, and both variants on Qwen2.5). These results indicate that general steering directions can transfer across instances while maintain- ing positive effects on CoT accuracy, providing a practical al- ternative that avoids instance-specific extraction at test time. In terms of average performance gain ∆ over the CoT base- line across models, Vanilla Steering achieves an improvement of 1.9%, SGS 1.4%, and TGS 2.5%, with TGS yielding the largest gain. Notably, TGS derives its steering vectors purely from textual data, without using spoken inputs during extrac- tion. This suggests that certain reasoning-relevant representa- tion directions present in the text modality can transfer to spo- ken tasks. Such cross-modal transfer implies that steering di- rections may capture modality-agnostic reasoning patterns, of- fering a lightweight way to enhance spoken reasoning without additional speech-specific training. 4.2. Hyperparameter Sensitivity of Steering Methods We analyze the hyperparameter sensitivity of the steering meth- ods on Voxtral, where all methods show clear improvements. Following Table 1, we vary the scaling factor α while fixing k, and vice versa, reporting average accuracy across benchmarks. Figure 4: Effect of external dataset size for SGS/TGS on Voxtral. As shown in Figure 3a, Vanilla Steering is highly sensitive to α: performance peaks at small values and degrades rapidly as α increases. Instance-specific steering vectors capture input- dependent representation shifts, which can be over-amplified under large scaling factors, leading to unstable predictions. In contrast, SGS and TGS remain comparatively stable across a wider range of α, likely because their aggregated steering di- rections induce smoother representation shifts. Figure3bshowsthatperformancevariesnon- monotonically with respect to the number of steered last layers.Increasing k does not consistently improve perfor- mance, and Vanilla Steering exhibits larger fluctuations than others.These results indicate that steering effectiveness depends on both scaling strength and layer position, with instance-specific steering being more sensitive to hyperparam- eter choices.Developing automatic strategies for selecting hyperparameters is a promising direction for future research. 4.3. Data Efficiency of SGS and TGS We analyze the data efficiency of SGS and TGS by varyingD s ext andD t ext (five-run average; Figure 4). For SGS, accuracy increases steadily with more spoken samples. Gains are largest at small scales (1–10 samples) and remain evident up to around 40 samples, after which perfor- mance largely saturates with diminishing returns. This suggests that estimating a reliable shared steering direction from speech requires a moderate amount of spoken data. In contrast, TGS remains relatively stable across data scales and reaches near-peak performance even with few textual sam- ples (e.g., 10), indicating robustness to dataset size. Because TGS does not rely on spoken data during extraction, it is more data-efficient and practical when spoken data is limited. We attribute this stability to the generally stronger and more consistent reasoning performance of LALMs in text-based set- tings [20, 53]. Since spoken reasoning is typically more chal- lenging [50, 51, 54], steering directions derived from text can be estimated more reliably under limited data. 5. Conclusion In this paper, we investigate inference-time model steering as a training-free approach to enhance Chain-of-Thought reasoning in large audio-language models. Across four LALMs and four spoken reasoning benchmarks, we show that representation- level intervention can consistently improve CoT performance. While instance-specific steering achieves strong gains, it is sen- sitive to hyperparameters and less stable. In contrast, gener- alized steering directions can be estimated from a moderate amount of auxiliary data and reused across inputs. Notably, text-derived steering provides stable improvements on speech- based tasks with only a few samples, highlighting cross-modal transferability and data efficiency. Overall, our results demon- strate the practical feasibility of model steering for inference- time CoT enhancement in LALMs. 6. Acknowledgement We acknowledge the computational and storage support pro- vided by the National Center for High-performance Comput- ing (NCHC) of the National Applied Research Laboratories (NARLabs) in Taiwan. 7. Generative AI Use Disclosure In this work, generative AI was used only to improve clarity and writing, without contributing to the research content. 8. References [1] A. H. Liu, A. Ehrenberg, A. Lo, C. Denoix, C. Barreau, G. Lam- ple, J.-M. Delignon, K. R. Chandu, P. von Platen, P. R. Mud- direddy et al., “Voxtral,” arXiv preprint arXiv:2507.13264, 2025. [2] A. Abouelenin, A. Ashfaq, A. Atkinson, H. Awadalla, N. Bach, J. Bao, A. Benhaim, M. Cai, V. Chaudhary, C. Chen et al., “Phi-4-mini technical report:Compact yet powerful multi- modal language models via mixture-of-loras,” arXiv preprint arXiv:2503.01743, 2025. [3] J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y. Fan, K. Dang et al., “Qwen2. 5-omni technical report,” arXiv preprint arXiv:2503.20215, 2025. [4] S. Ghosh, A. Goel, J. Kim, S. Kumar, Z. Kong, S. gil Lee, C.-H. H. Yang, R. Duraiswami, D. Manocha, R. Valle, and B. Catanzaro, “Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,” in The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Online]. Available:https://openreview.net/forum?id= FjByDpDVIO [5] K.-H. Lu, Z. Chen, S.-W. Fu, C.-H. H. Yang, S.-F. Huang, C.- K. Yang, C.-E. Yu, C.-W. Chen, W.-C. Chen, C.-y. Huang et al., “Desta2. 5-audio: Toward general-purpose large audio language model with self-generated cross-modal alignment,” arXiv preprint arXiv:2507.02768, 2025. [6] C.-Y. Kuan et al., “Speech-copilot: Leveraging large language models for speech processing via task decomposition, modular- ization, and program generation,” in 2024 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2024, p. 1060–1067. [7] C.-K. Yang, Y.-K. Fu, C.-A. Li, Y.-C. Lin, Y.-X. Lin, W.-C. Chen, H. L. Chung, C.-Y. Kuan, W.-P. Huang, K.-H. Lu et al., “Building a taiwanese mandarin spoken language model: A first attempt,” arXiv preprint arXiv:2411.07111, 2024. [8] Y.-X. Lin, C.-K. Yang, W.-C. Chen, C.-A. Li, C.-y. Huang, X. Chen, and H.-y. Lee, “A preliminary exploration with gpt-4o voice mode,” arXiv preprint arXiv:2502.09940, 2025. [9] A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford et al., “Gpt-4o system card,” arXiv preprint arXiv:2410.21276, 2024. [10] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al- Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan et al., “The llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024. [11] A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv et al., “Qwen3 technical report,” arXiv preprint arXiv:2505.09388, 2025. [12] C.-K. Yang, N. Ho, Y.-J. Lee, and H.-y. Lee, “Audiolens: A closer look at auditory attribute perception of large audio-language mod- els,” arXiv preprint arXiv:2506.05140, 2025. [13] C.-K. Yang, Y.-T. Piao, T.-W. Hsu, S.-W. Fu, Z. Chen, K.-H. Lu, S.-F. Huang, C.-H. H. Yang, Y.-C. F. Wang, Y.-N. Chen et al., “Sake: Towards editing auditory attribute knowledge of large audio-language models,” arXiv preprint arXiv:2510.16917, 2025. [14] C.-y. Huang et al., “Dynamic-SUPERB phase-2: A collabora- tively expanding benchmark for measuring the capabilities of spo- ken language models with 180 tasks,” in The Thirteenth Interna- tional Conference on Learning Representations, 2025. [15] S. Arora, K.-W. Chang, C.-M. Chien, Y. Peng, H. Wu, Y. Adi, E. Dupoux, H. yi Lee, K. Livescu, and S. Watanabe, “On the landscape of spoken language models: A comprehensive survey,” Transactions on Machine Learning Research, 2025. [Online]. Available: https://openreview.net/forum?id=BvxaP3sVbA [16] C.-K. Yang, N. S. Ho, and H.-y. Lee, “Towards holistic evaluation of large audio-language models: A comprehensive survey,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng, Eds.Suzhou, China: Association for Computational Linguistics, Nov. 2025, p. 10 144–10 170. [Online]. Available:https://aclanthology.org/ 2025.emnlp-main.514/ [17] S. Wang, Z. Jin, C. Tang, Q. Li, B. Li, C. Chen, Y. Hu, W. Yu, Y. Li, J. Zhuang et al., “Towards general auditory intelligence: Large multimodal models for machine listening and speaking,” arXiv preprint arXiv:2511.01299, 2025. [18] S. Sakshi et al., “MMAU: A massive multi-task audio understand- ing and reasoning benchmark,” in The Thirteenth International Conference on Learning Representations, 2025. [19] S. Kumar, ˇ S. Sedl ́ a ˇ cek, V. Lokegaonkar, F. L ́ opez, W. Yu, N. Anand, H. Ryu, L. Chen, M. Pli ˇ cka, M. Hlav ́ a ˇ cek et al., “Mmau-pro: A challenging and comprehensive benchmark for holistic evaluation of audio general intelligence,” arXiv preprint arXiv:2508.13992, 2025. [20] C.-K. Yang, N. Ho, Y.-T. Piao, and H. yi Lee, “SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information,” in Interspeech 2025, 2025, p. 1788–1792. [21] T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” in Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., vol. 35. Curran Associates, Inc., 2022, p. 22 199–22 213. [Online]. Available: https://proceedings.neurips.c/paper files/paper/2022/ file/8b0d291acd4acf06ef112099c16f326-Paper-Conference.pdf [22] J. Wei, X. Wang, D. Schuurmans, M. Bosma, b. ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., vol. 35. Curran Associates, Inc., 2022, p. 24 824–24 837. [Online]. Available: https://proceedings.neurips.c/paper files/paper/2022/ file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf [23] Z. Ma, Z. Chen, Y. Wang, E. S. Chng, and X. Chen, “Audio- cot: Exploring chain-of-thought reasoning in large audio language model,” arXiv preprint arXiv:2501.07246, 2025. [24] Z. Xie, M. Lin, Z. Liu, P. Wu, S. Yan, and C. Miao, “Audio- reasoner: Improving reasoning capability in large audio language models,” arXiv preprint arXiv:2503.02318, 2025. [25] S. Wu, C. Li, W. Wang, H. Zhang, H. Wang, M. Yu, and D. Yu, “Audio-thinker: Guiding audio language model when and how to think via reinforcement learning,” arXiv preprint arXiv:2508.08039, 2025. [26] N. Subramani, N. Suresh, and M. Peters, “Extracting latent steering vectors from pretrained language models,” in Findings of the Association for Computational Linguistics: ACL 2022, S. Muresan, P. Nakov, and A. Villavicencio, Eds.Dublin, Ireland: Association for Computational Linguistics, May 2022, p. 566–581. [Online]. Available: https://aclanthology.org/2022. findings-acl.48/ [27] T.-M. Pai, J.-I. Wang, L.-C. Lu, S.-H. Sun, H.-Y. Lee, and K.- W. Chang, “Billy: Steering large language models via merg- ing persona vectors for creative generation,” arXiv preprint arXiv:2510.10157, 2025. [28] A. M. Turner, L. Thiergart, G. Leech, D. Udell, J. J. Vazquez, U. Mini, and M. MacDiarmid, “Steering language models with activation engineering,” 2025. [Online]. Available: https://openreview.net/forum?id=2XBPdPIcFK [29] A. Arditi, O. Obeso, A. Syed, D. Paleka, N. Panickssery, W. Gurnee, and N. Nanda, “Refusal in language models is medi- ated by a single direction,” in Advances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, Eds., vol. 37. Curran Associates, Inc., 2024, p. 136 037–136 083. [Online]. Available: https://proceedings.neurips.c/paper files/paper/2024/ file/f545448535dfde4f9786555403ab7c49-Paper-Conference.pdf [30] C. Venhoff, I. Arcuschin, P. Torr, A. Conmy, and N. Nanda, “Understanding reasoning in thinking language models via steering vectors,” in Workshop on Reasoning and Planning for Large Language Models, 2025. [Online]. Available: https: //openreview.net/forum?id=OwhVWNOBcz [31] R. Zhu, Y. Wang, T. Jiang, J. Liang, and T. Wang, “Self-improving model steering,” arXiv preprint arXiv:2507.08967, 2025. [32] V. Sinii, A. Gorbatovski, A. Cherepanov, B. Shaposhnikov, N. Balagansky, and D. Gavrilov, “Steering LLM reasoning through bias-only adaptation,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Pro- cessing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng, Eds.Suzhou, China: Association for Computational Linguistics, Nov. 2025, p. 9202–9211. [Online]. Available: https://aclanthology.org/2025.emnlp-main.467/ [33] T.-E. Lin, K.-Y. Lee, and H.-Y. Lee, “Adaptive vector steering: A training-free, layer-wise intervention for hallucination miti- gation in large audio and multimodal models,” arXiv preprint arXiv:2510.12851, 2025. [34] W. Lin, J. Li, H. Xiong, and L. Liu, “Sarsteer: Safeguarding large audio language models via safe-ablated refusal steering,” arXiv preprint arXiv:2510.17633, 2025. [35] X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” in The Eleventh International Conference on Learning Representations, 2023. [Online]. Available:https://openreview.net/forum?id= 1PL1NIMMrw [36] M. Valentino, G. Kim, D. Dalal, Z. Zhao, and A. Freitas, “Mitigat- ing content effects on reasoning in language models through fine- grained activation steering,” arXiv preprint arXiv:2505.12189, 2025. [37] Y. Huang, H. Chen, S. Ruan, Y. Zhang, X. Wei, and Y. Dong, “Mitigating overthinking in large reasoning models via manifold steering,” in The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Online]. Available: https://openreview.net/forum?id=49Rc51iCso [38] A. Wang, D. Shu, Y. Wang, Y. Ma, and M. Du, “Improving LLM reasoning through interpretable role-playing steering,” in Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng, Eds. Suzhou, China: Association for Computational Linguistics, Nov. 2025, p. 731–751. [Online]. Available: https://aclanthology.org/2025.findings-emnlp.39/ [39] Y. Tang, K. Zhou, Y. Min, W. X. Zhao, J. Sha, Z. Sheng, and S. Wang, “Enhancing chain-of-thought reasoning via neuron activation differential analysis,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng, Eds.Suzhou, China: Association for Computational Linguistics, Nov. 2025, p. 16 151–16 159. [Online]. Available: https://aclanthology.org/2025.emnlp-main.817/ [40] A. Stolfo, V. Balachandran, S. Yousefi, E. Horvitz, and B. Nushi, “Improving instruction-following in language models through activation steering,” in The Thirteenth International Conference on Learning Representations, 2025. [Online]. Available: https://openreview.net/forum?id=wozhdnRCtw [41] J. Braun, C. Eickhoff, D. Krueger, S. A. Bahrainian, and D. Krasheninnikov, “Understanding (un)reliability of steering vectors in language models,” in ICLR 2025 Workshop on Foundation Models in the Wild, 2025. [Online]. Available: https://openreview.net/forum?id=qGCp2AYosf [42] N. Rimsky, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. Turner, “Steering llama 2 via contrastive activation addition,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L.-W. Ku, A. Martins, and V. Srikumar, Eds.Bangkok, Thailand:Association for Computational Linguistics, Aug. 2024, p. 15 504–15 522. [Online]. Available: https://aclanthology.org/2024.acl-long.828/ [43] N. Belrose, “Diff-in-means concept editing is worst-case optimal: Explaining a result by sam marks and max tegmark,” https://blog. eleuther.ai/diff-in-means/, December 2023. [44] S. Marks and M. Tegmark, “The geometry of truth: Emergent linear structure in large language model representations of true/false datasets,” in First Conference on Language Modeling, 2024. [Online]. Available:https://openreview.net/forum?id= aajyHYjjsk [45] C. Tigges, O. J. Hollinsworth, A. Geiger, and N. Nanda, “Lin- ear representations of sentiment in large language models,” arXiv preprint arXiv:2310.15154, 2023. [46] K.-H. Lu, C.-Y. Kuan, and H. yi Lee, “Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models,” in Interspeech 2025, 2025, p. 2078–2082. [47] [ByteDance-Seed],“Beyondaime:Advancing math rea- soning evaluation beyond high school olympiads,” [https: //huggingface.co/datasets/ByteDance-Seed/BeyondAIME](https: //huggingface.co/datasets/ByteDance-Seed/BeyondAIME), 2025. [48] S. Zhou, Y. Zhou, Y. He, X. Zhou, J. Wang, W. Deng, and J. Shu, “Indextts2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech,” arXiv preprint arXiv:2506.21619, 2025. [49] K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano et al., “Train- ing verifiers to solve math word problems,” arXiv preprint arXiv:2110.14168, 2021. [50] W. Yang, Y. Li, Y. Wei, M. Fang, and L. Chen, “Speechr: A bench- mark for speech reasoning in large audio-language models,” arXiv preprint arXiv:2508.02018, 2025. [51] W. Cui, X. Jiao, Z. Meng, and I. King, “VoxEval: Benchmarking the knowledge understanding capabilities of end-to-end spoken language models,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds.Vienna, Austria: Association for Computational Linguistics, Jul. 2025, p. 16 735–16 753. [Online]. Available: https://aclanthology.org/2025.acl-long.818/ [52] W. Yang, Y. Li, Y. Wei, M. Fang, and L. Chen, “Speechr: A bench- mark for speech reasoning in large audio-language models,” arXiv preprint arXiv:2508.02018, 2025. [53] C.-A. Li, T.-H. Lin, and H.-y. Lee, “When silence matters: The impact of irrelevant audio on text reasoning in large audio- language models,” arXiv preprint arXiv:2510.00626, 2025. [54] C.-Y. Hsiao, K.-H. Lu, K.-W. Chang, C.-K. Yang, W.-C. Chen, and H. yi Lee, “Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models,” in Interspeech 2025, 2025, p. 3234–3238.