Paper deep dive
EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$
Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/28/2026, 3:14:09 AM
Summary
The paper introduces EXAM², a benchmark for multilingual and multimodal audio understanding that spans six languages (English, German, Spanish, Japanese, Malay, Chinese) and multiple audio modalities (speech, sound, music, mixed) combined with visual images. It comprises 5,667 multiple-choice questions and 22,614 image instances. The authors evaluate state-of-the-art Large Audio Language Models (LALMs) and Multimodal LLMs, revealing performance gaps. They propose Gemma3n-EXAM², a lightweight fusion model fine-tuned using OmniLoRA, which achieves significant improvements in multilingual and multimodal settings.
Entities (18)
Relation Signals (15)
EXAM² → supportslanguages → Japanese
confidence 95% · EXAM² ... spanning six languages: English, German, Spanish, Japanese, Malay, and Chinese.
EXAM² → supportslanguages → Spanish
confidence 95% · EXAM² ... spanning six languages: English, German, Spanish, Japanese, Malay, and Chinese.
EXAM² → supportslanguages → German
confidence 95% · EXAM² ... spanning six languages: English, German, Spanish, Japanese, Malay, and Chinese.
EXAM² → supportslanguages → English
confidence 95% · EXAM² ... spanning six languages: English, German, Spanish, Japanese, Malay, and Chinese.
EXAM² → supportslanguages → Malay
confidence 95% · EXAM² ... spanning six languages: English, German, Spanish, Japanese, Malay, and Chinese.
EXAM² → supportslanguages → Chinese
confidence 95% · EXAM² ... spanning six languages: English, German, Spanish, Japanese, Malay, and Chinese.
Gemma3n-EXAM² → isfinetunedfrom → Gemma3n-E4B
confidence 92% · we propose Gemma3n-EXAM², a lightweight fusion-model fine-tuned on EXAM²-train ... We build on Gemma3n-E4B
EXAM² → supportsmodalities → Sound
confidence 90% · multiple modalities, including speech, sound, music, mixed-audio settings
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent large audio language models (LALMs) have achieved impressive progress in audio understanding. However, existing evaluations remain largely constrained to English and narrow audio domains. Prior benchmarks typically focus on a single audio modality, i.e., speech, sound, or music, limiting the systematic investigation into how these models generalize across diverse visual scenarios. In this paper, we introduce EXAM$^2$, a benchmark for multilingual and multimodal audio understanding spanning six languages and multiple modalities, including speech, sound, music, mixed-audio settings, and visual images. By incorporating visual information alongside heterogeneous audio inputs, EXAM$^2$ enables more realistic evaluation of scene-aware audio reasoning and cross-modal comprehension. EXAM$^2$ comprises $5,667$ multiple-choice questions, $22,614$ image instances, and $135,684$ multilingual translations. We evaluate state-of-the-art open-source and proprietary LALMs as well as multimodal LLMs, revealing substantial performance gaps in multilingual and cross-modal understanding. Furthermore, we propose Gemma3n-EXAM$^2$, a lightweight fusion-model fine-tuned on EXAM$^2$-train, achieves up to $12.4\%$ improvement in multilingual settings and $21.7\%$ gains in multimodal evaluation over a strong baseline. Empirical results establish EXAM$^2$ as a challenging benchmark and pioneer future multilingual and multimodal audio intelligence research.
Tags
Links
- Source: https://arxiv.org/abs/2608.23758v2
- Canonical: https://arxiv.org/abs/2608.23758v2
Trouble viewing inline? Open PDF directly →
Full Text
58,110 characters extracted from source content.
Expand or collapse full text
EXAM 2 :ExtendingAudio Understanding inMultilingual andMultimodal Analysis Jiawen Wang 1 Xiaoxue Gao 2 Zi Haur Pang 3 Nancy F. Chen 4 2 School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China 1 LMU Munich, Germany 3 Kyoto University, Japan 4 A*STAR, Singapore jiawen.wang@campus.lmu.de Abstract Recent large audio language models (LALMs) have achieved impressive progress in audio un- derstanding. However, existing evaluations re- main largely constrained to English and narrow audio domains. Prior benchmarks typically fo- cus on a single audio modality, i.e., speech, sound, or music, limiting the systematic in- vestigation into how these models generalize across diverse visual scenarios. In this paper, we introduce EXAM 2 , a benchmark for mul- tilingual and multimodal audio understanding spanning six languages and multiple modali- ties, including speech, sound, music, mixed- audio settings, and visual images. By incorpo- rating visual information alongside heteroge- neous audio inputs, EXAM 2 enables more re- alistic evaluation of scene-aware audio reason- ing and cross-modal comprehension. EXAM 2 comprises5, 667multiple-choice questions, 22, 614image instances, and135, 684multi- lingual translations. We evaluate state-of-the- art open-source and proprietary LALMs as well as multimodal LLMs, revealing substan- tial performance gaps in multilingual and cross- modal understanding. Furthermore, we pro- pose Gemma3n-EXAM 2 , a lightweight fusion- model fine-tuned on EXAM 2 -train, achieves up to12.4%improvement in multilingual settings and21.7%gains in multimodal evaluation over a strong baseline. 1 1 Introduction Recent advancement of large audio language mod- els (LALMs) (Deshmukh et al., 2023; Chu et al., 2024; Kong et al., 2024; Ghosh et al., 2025, 2026) has significantly expanded machine capabilities for understanding, reasoning over, and interacting with auditory signals (Wang et al., 2025b; Pang et al., 2026). To assess these capabilities, multiple choice question (MCQ) tasks (He et al., 2025) 1 All code, data, and models are available onhttps:// github.com/werywjw/EXAM-2. are being utilized across distinct audio modali- ties, including speech (Zhao et al., 2024; Huang et al., 2024), sound (Lipping et al., 2022; Ghosh et al., 2024), and music (Melechovsky et al., 2024; Weck et al., 2024), to evaluate how models inter- pret and reason about audio content under various context and scenarios. Consequently, a wide ar- ray of benchmarks have emerged, including Au- dioBench (Wang et al., 2025a), MMAU (Sakshi et al., 2025), and MMAR (Ma et al., 2025). Yet, these benchmarks remain primarily audio-only and rarely investigate interactions between heteroge- neous audio sources and visual grounding (Yue et al., 2024, 2025; Wang et al., 2025c; Chen et al., 2025). Beyond cross-modal limitations, restricted lan- guage coverage remains a critical bottleneck in current audio understanding evaluations. How- ever, the semantic interpretation of audio content is frequently intertwined with linguistic and cul- tural contexts (Werner, 2023); for instance, iden- tical lyrical slang can convey different meanings in diverse languages (Werner, 2023). Yet, existing audio-understanding benchmarks largely rely on English-centric question-answer settings (Drossos et al., 2020; Yang et al., 2024; Sakshi et al., 2025), overlooking models’ capabilities to navigate these shifting contexts. This discrepancy underscores the necessity for multilingual question-answer bench- marks to facilitate a broader assessment of audio comprehension across diverse language settings. To address these gaps, we introduce EXAM 2 , a benchmark for multilingual and multimodal au- dio understanding under a unifiedMCQevaluation framework (cf. Figure 1). EXAM 2 extends conven- tional audio-only benchmarks by jointly combin- ing speech, sound, music, mixed-audio scenarios, and visual representations across six languages: English, German, Spanish, Japanese, Malay, and Chinese. Beyond measuring recognition accuracy, our benchmark probes multilingual reasoning, se- arXiv:2608.23758v2 [cs.SD] 27 Aug 2026 ! ! ! ! Original English Question & Choices: What is the relationship between the two individuals in the conversation? ! ! immigration officer-traveler artist-art collector driver-passenger fire marshal-event planner German Translation: Welche Beziehung besteht zwischen den beiden Personen im Gespräch? Einwanderungsbeamter–Reisender Künstler–Kunstsammler Fahrer–Passagier Brandschutzbeauftragter–Veranstaltungsplaner Spanish Translation: ¿Cuál es la relación entre las dos personas en la conversación? oficial de inmigración-viajero artista-coleccionista de arte conductor-pasajero jefe de bomberos-organizador de eventos Japanese Translation: 会話に登場する二人はどのような関係ですか? 入国審査官-旅行者 アーティスト-美術収集家 運転手-乗客 消防監督官-イベントプランナー Malay Translation: Apakah hubungan antara kedua-dua individu dalam perbualan itu? pegawai imigresen - pengembara artis - pengumpul seni pemandu - penumpang pegawai pengawas kebakaran - perancang acara Chinese Translation: 对话中两个人之间的关系是什么? 移民官-旅客艺术家-艺术收藏家 司机-乘客 消防官-活动策划人 EXAM example with Multilingual & Multimodal Extension 2 Valid. ✅ Figure 1: Overview of a multimodal and multilingual example from our EXAM 2 Benchmark. mantic grounding, and cross-modal comprehension in contemporaryLALMs, multimodal large lan- guage models (MLLMs), and cascaded systems via large language models (LLMs). Finally, our proposed lightweight Gemma3n-EXAM 2 model achieves up to+21.65%average multilingual im- provement compared to the baseline, demonstrat- ing the model’s effectiveness for advancing multi- lingual and multimodal audio MCQ performance. Our contributions mainly include: (1) We pro- pose the EXAM 2 benchmark, to our best knowl- edge, the first to address multilingual and mul- timodal audio understanding with5, 667 MCQs, 22, 614image instances, and135, 684multilin- gual translated instances; (2) We introduce Om- niLoRA, a straightforward omni-language uniform tuning based on LoRA, extensive experiments show the effectiveness of our method; (3) We present our multilingual Gemma3n-EXAM 2 , which is a lightweight fusion approach with significant im- provements across languages to foster future re- search in multilingual and multimodal domains. 2 Related Work Large Audio Language Models. Early works such as AudioCLIP (Guzhov et al., 2022) and CLAP (Elizalde et al., 2023) focused on learning shared audio-text representations for retrieval and captioning tasks. Building on these foundations, LALMs such as Qwen2-Audio (Chu et al., 2024), Audio Flamingo (Kong et al., 2024; Ghosh et al., 2025, 2026), and SALMONN (Tang et al., 2024) integrate audio encoders with LLMs to support in- struction following and complex audio understand- ing. In parallel, general-purpose omni-models such as Qwen-omni (Xu et al., 2025; Team, 2026), Ming- omni (AI et al., 2025) and Baichuan-omni (Li et al., 2025) have shown strong cross-modal generaliza- tion despite not being specifically designed for au- dio tasks. While prior benchmarks mainly focus on evaluatingLALMs, our work provides a com- prehensive benchmark acrossLALMs,MLLMs, and cascadedLLMs that first transcribe audio be- fore performing reasoning with (large) language or reasoning models. Multilingual and Multimodal Audio Under- standing Benchmarks. Previous benchmarks have advanced the evaluation ofLALMs across speech, sound, and music understanding. Domain specific benchmarks such as LibriSQA (Zhao et al., 2024) and Dynamic-SUPERB (Huang et al., 2024) for speech-related tasks, Clotho-AQA (Lipping et al., 2022), CompA (Ghosh et al., 2024) for en- vironmental sound reasoning (Iyer et al., 2026), and MusicBench (Melechovsky et al., 2024) and MuChoMusic (Weck et al., 2024) for music un- Benchmark TextAudioVision Language Tr / CaSp / So / Mu / MixImageENDEESJAMSZH Clotho✘ /✔ /✔ /✔ /✘✔✘ Clotho-AQA✘ /✘ /✔ /✘ /✘✔✘ LibriSQA✘ /✘✔ /✘ /✘ /✘ ✔✘ Dynamic-SUPERB✘ /✘✔ /✘ /✘ /✘✔✘ MusicBench✘ /✘ /✘ /✔ /✘✔✘ MuChoMusic✘ /✘ /✘ /✔ /✘✔✘ AIR-Bench✘ /✘✔ /✔ /✔ /✔✘✔✘ AudioSetCaps✘ /✔ /✔ /✔ /✘✔✘ AudioMCQ✘ /✘✔ /✔ /✔ /✘✔✘ MMAR✘ /✘✔ /✔ /✔ /✔✘✔✘✔✘✔ MMAU✘ /✘✔ /✔ /✔ /✘✔✘ MMAU-Pro✘ /✘✔ /✔ /✔ /✘✔✘ MMMU✘ /✘ /✘ /✘ /✘✔✘ MMAU-Pro✘ /✘ /✘ /✘ /✘✔✘ EXAM 2 (Ours)✔ /✔ /✔ /✔ /✔ Table 1: The list of existing multimodal and multilin- gual datasets show the novelty of our EXAM 2 bench- mark. Tr, Ca denotes Transcript and Caption, respec- tively. Sp / So / Mu / Mix stands for Speech, Sound, Music, and their Mix. derstanding. Broader benchmarks such as Au- dioBench (Wang et al., 2025a), AIR-Bench (Yang et al., 2024), MMAU (Sakshi et al., 2025), MMAU- Pro (Kumar et al., 2026), and MMAR (Ma et al., 2025) further combine multiple audio domains and evaluate advanced auditory reasoning capabili- ties. However, existing benchmarks remain largely audio-only, with limited support for multimodal grounding or visual representations in audio under- standing tasks. In addition, multilingual evaluation is still underexplored. Prior multilingual bench- marks such as BUFFET (Asai et al., 2024) and MEGA (Ahuja et al., 2023) focus on text-based LLMs rather than theMLLMs, while audio bench- mark MMAR provides very little multilingual au- dio and entirely lack multilingual question-and- answer pairs. 3 EXAM 2 Benchmark 3.1 Overview Our EXAM 2 introduces multilingual and multi- modal audio understanding with visual representa- tions, supporting six languages (DE, ES, JA, MS, ZH, and EN) to enable richer cross-modal and cross-lingual reasoning (cf. Table 1). Dataset statis- tics for our proposed EXAM 2 benchmark is given in Table 2. The benchmark consists of two main subsets: EXAM 2 -train and EXAM 2 -test, with a total of5, 667audio instances,135, 684multilin- gual language instances, and22, 614visual answer choices. EXAM 2 -train, recreated and filtered from MMAR and Clotho, includes4, 669questions cov- ering four audio domains (speech, sound, music, and mix) with18, 676visual choices. EXAM 2 - test, restructured and annotated from MMAU-test- BenchmarkStatisticsNumber EXAM 2 -test Question998 Audio Domains3 Domain CategoriesSpeech / Sound / Music Category Distribution333:333:332 (1:1:1) Visual Choices3,938 LanguagesEN / DE / ES / JA / MS / ZH Multilingual Instances23,628 EXAM 2 -train Question4669 Audio Domains4 Domain CategoriesSpeech / Sound / Music / Mix Category Distribution396:3737:243:293 Visual Choices18,676 LanguagesEN / DE / ES / JA / MS / ZH Multilingual Instances112,056 EXAM 2 Question5667 Audio Domains4 Domain CategoriesSpeech / Sound / Music / Mix Category Distribution729:4070:575:293 Visual Choices22,614 LanguagesEN / DE / ES / JA / MS / ZH Multilingual Instances135,684 Average Question Length10.7 words Average Option Length3.2 words Average Audio Duration12.9 sec Table 2: Dataset statistics of our EXAM 2 benchmark. mini, contains998questions across three audio domains (speech, sound, and music) with a bal- anced category distribution and a total of3, 938 visual choices. 3.2 Dataset Construction Figure 2 illustrates the construction pipeline of our EXAM 2 benchmark with seven key steps. Source Selection.We begin by collecting diverse audio corpora spanning speech, music, and envi- ronmental sounds, prioritizing real-world record- ings over synthetic data to ensure ecological va- lidity. Specifically, we curate four representative datasets: MMAU-test-mini (Sakshi et al., 2025), MMAR (Ma et al., 2025), Clotho (Drossos et al., 2020), and AudioMCQ-Clotho (He et al., 2025) un- derCreative Commonslicense to ensure a strong foundation for task development. Quality Control and Filtering.To ensure bench- mark quality, all collected audio samples, textual questions, choices, and answers undergo manual inspection by the authors. For MMAU-test-mini, we remove two music-category samples that fail to satisfy quality standards and eliminate dupli- cate answer choices within the same question, re- sulting in a final test set of998instances with 3, 938answer choices. For MMAR, we retrieve audio samples from the provided URLs and dis- card129out of1, 000instances due to inaccessible or broken links, yielding871valid audio instances with3, 484answer choices. For the Clotho bench- EXAM Benchmark Construction Pipeline 2 Source Selection Quality Control & Filtering Image Generation Language Translation Human Validation Expert Review Figure 2: EXAM 2 Benchmark Construction Pipeline. mark, we derive multiple-choice questions from AudioMCQ-Clotho, obtaining3, 798audios with 15, 192 choices. Image Generation. To enable multimodal eval- uation, we generate visual counterparts for all an- swer choices in EXAM 2 . For the smaller subsets, EXAM 2 -MMAU and EXAM 2 -MMAR, we employ Stable Diffusion (Rombach et al., 2022) for image generation underCreativeML Open RAIL-Mli- cense. For EXAM 2 -Clotho, visual answer choices are generated using GPT-image-2. 2 Human Validation. All textual questions and generated visual candidates are manually reviewed by the experts to ensure semantic relevance, clar- ity, and overall quality. When generated images are judged to be misleading, ambiguous, or insuf- ficiently aligned with the corresponding answer choice, we manually curate alternative visual rep- resentations. Representative examples across cate- gories are provided in Section A.2.2. Language Translation.To support multilingual evaluation, we translate the original English ques- tions and answer choices into five target languages: German (DE), Spanish (ES), Japanese (JA), Malay (MS), and Chinese (ZH). We employ GPT-5-mini 3 for translation, using carefully designed prompts (See Section A.3) to preserve semantic fidelity and task consistency across languages. Domain- specific terminology related to audio understanding is further reviewed and manually refined where nec- essary to ensure accurate translation and conceptual equivalence. Expert Review.Translated questions and answer choices are reviewed by either native speakers of 2 https://openai.com/index/ introducing-chatgpt-images-2-0/ 3 https://openai.com/index/introducing-gpt-5/ the target languages or experienced language users with more than five years of advanced proficiency. The review process focuses on semantic correct- ness, linguistic fluency, and preservation of task intent across languages, ensuring high-quality mul- tilingual question-answer pairs across all bench- mark languages. 4 Methodology 4.1 Task Formulation We address multimodalMCQover audio and visual inputs. Each instance is a tuple(a, V, q, C, y ∗ ), wherea∈ R T is a raw audio waveform ofTsam- ples,V = v 1 ,...,v K is a set ofKcandidate images (one per choice),qis a natural-language question,C =c 1 ,...,c K is the set of candidate answers, andy ∗ ∈ Cis the ground-truth answer. The model must identifyy ∗ by jointly reasoning over both modalities. 4.2 Multimodal Input Representation We build on Gemma3n-E4B 4 , aMLLMthat en- codes audio, images, and text within a unified trans- former. Audio. The waveformais resampled to 16 kHz and converted to a spectrogramS∈ R F×L (Ffre- quency bins,Lframes). A convolutional audio en- coder projects S into a sequence of d-dimensional hidden states H (a) ∈ R L ′ ×d , which are prepended to the token sequence as soft audio tokens. Images.For each choicec k , a corresponding im- agev k ∈ R H×W×3 is encoded into patch embed- dingsH (v k ) ∈ R N p ×d , whereN p is the number of patches. The image embeddings for allKchoices 4 https://huggingface.co/google/ gemma-3n-E4B-it are concatenated and interleaved with textual to- kens. Multimodal Fusion. The combined input se- quence fed to the language-model backbone is X = [ H (v 1 ) ,..., H (v K ) , H (a) , e q , e C ],(1) wheree q ande C are the token embeddings of the question and choice text, respectively. No modality- specific fusion layer is added; cross-modal align- ment is learned entirely through self-attention. 4.3 Parameter-Efficient Fine-Tuning via LoRA Full fine-tuning of all model parametersθis compu- tationally prohibitive. Following Hu et al. (2022), we propose to freezeθand inject trainable low-rank update matrices into the attention and feed-forward layers. We fine-tune the model with a cross-entropy loss restricted to the answer token. Given the mul- timodal inputX(Eq. 1), the model is trained to predict the single-digit indexy ∗ ∈ 1,...,K identifying the correct choice: L(θ LoRA ) =− logP θ (y ∗ | X).(2) After training, the LoRA adapters are merged back into the base weights and saved as a single self-contained model for inference. To support multilingual evaluation, we propose omni-language uniform tuning (OmniLoRA), which applies LoRA to backbone model by uniformly sampling one of six languagesℓ ∈ EN, DE, ES, JA, MS, ZH for each training instance per epoch: L Omni (θ LoRA ) =− E ℓ∼L h logP θ y ∗ | X (ℓ) i , (3) whereX (ℓ) denotes the fused multimodal sequence (Eq. 1) with text rendered in language ℓ. 5 Evaluation 5.1 Experimental Setup We evaluate a diverse set of closed- and open- source models on proposed EXAM 2 benchmark. Our results can be easily reproduced using the provided codebase and evaluation scripts. Closed- source models and Phi-4-multimodal-instruct are conducted using Azure OpenAI API, while the re- maining open-source models are evaluated using Hugging Face library. The supervised fine-tuning on our EXAM 2 -train allocates four NVIDIA A40 GPUs. For all models and languages, we use the same set of prompts, detailed in Appendix A.3, across audio domains (speech, sound, and music) to ensure a fair comparison. 5.2 Baseline Closed- and open-source models are evaluated on the EXAM 2 -test, in zero-shot settings. For closed- source models, we evaluate GPT-4o-audio 5 and GPT-4o-mini-audio. GPT-4o 6 is leveraged as a MLLMwith image input, while GPT-5-mini 7 and GPT-5.2 8 are evaluated as text-only models to as- sess the impact of multimodal grounding and mul- tilingual competence. Phi-4-multimodal-instruct variant 9 , along with other models such as Qwen- 2.5-omni 10 and Gemma-3n-E4B 11 , are evaluated as open-source alternatives. 5.3 Results and Analysis Table 3 summarizes the average performance on the EXAM 2 benchmark across three audio domains (speech, sound, and music) and six languages (DE, EN, ES, JA, MS, and ZH). Among closed-source models, GPT-4o-audio achieves the strongest over- all performance with an average score of 73.83%, consistently leading in Speech, Sound, and Mu- sic. Focusing on open-source models, our pro- posed Gemma3n-EXAM 2 obtains the best over- all average of 61.24% (with significant improve- ment over Gemma3n, by paired t-test,p < 0.05), and achieves the highest scores in both Speech and Music with 64.31% and 60.14%, respectively. However, Gemma3n † remains stronger in Sound, achieving 67.12% compared with 59.26% for Gemma3n-EXAM 2 , indicating that our model is particularly effective for speech- and music-related understanding while leaving room for further im- provement on sound-centric reasoning. Meanwhile, Table 4 presents our results across all six languages. Overall, our proposed Gemma3n- EXAM 2 model, compared to baseline with the same modalities input, demonstrates competitive performance across all languages and audio cate- 5 https://developers.openai.com/api/docs/ models/gpt-4o-audio-preview 6 https://developers.openai.com/api/docs/ models/gpt-4o 7 https://openai.com/gpt-5/ 8 https://openai.com/index/ introducing-gpt-5-2/ 9 https://huggingface.co/microsoft/ Phi-4-multimodal-instruct 10 https://huggingface.co/Qwen/Qwen2.5-Omni-7B 11 https://huggingface.co/google/ gemma-3n-E4B-it Model Input ModalityAverage (%) ImageAudioSpeech Sound Music Avg. Closed-source GPT-4o-mini-audio✘✔73.6761.01 57.13 63.94 GPT-4o-audio✘✔ 77.9373.17 70.38 73.83 GPT-4o † ✔✓ 66.6769.17 57.13 64.32 GPT-5-mini † ✘✓ 71.8251.25 54.29 59.12 GPT-5-2 † ✘✓71.4750.25 50.10 57.27 Open-source Phi-4-multimodal-instruct✘✔59.9154.47 55.02 56.47 Phi-4-multimodal-instruct † ✘✓ 47.0560.81 56.21 54.69 Phi-4-multimodal-instruct † ✔✓51.3059.91 44.08 51.76 Qwen-2.5-omni † ✘✓ 53.6064.17 56.5258.10 Qwen-2.5-omni † ✔✓53.2561.11 55.37 56.58 Qwen-2.5-omni✘✔ 39.6449.90 32.58 40.71 Qwen-2.5-omni✔52.2061.61 54.97 56.26 Gemma3n † ✘✓ 58.4167.12 54.22 59.91 Gemma3n † ✔✓57.1666.3753.46 59.00 Gemma3n✔✘42.1448.70 46.64 45.83 Gemma3n✘✔41.9947.35 46.99 45.44 Gemma3n✔42.8443.04 48.34 44.74 Gemma3n-EXAM 2 (Ours)✔64.3159.26 60.14 61.24 Table 3: Average multilingual and multimodal evalua- tion results across the six languages over Speech, Sound, and Music. Bold andunderlineindicate the best and second best performance for each category among open- source models only. † shows cascaded input of transcript and/or caption.✔and✓are marked to differentiate the original audio waveform and textual input, respectively. ✘ marks the unexploited modality. gories, with a notable increase of26.43%in the speech domain for German and+21.65%accuracy gain for all languages on average. A global im- provement of11.91%and13.15%in the sound and music domain respectively is observed, as both sound and music benefit from the multi- modal grounding provided by the visual answer choices. By utilizing our OmniLoRA approach on Gemma3n-E4B with our multilingual and mul- timodal training data, we can significantly boost the performance of the model across all languages and audio domains, demonstrating the high quality of our benchmark for improving multilingual and multimodal audio understanding. Image modality helps audio understanding. For all languages, we notice a significant improve- ment by incorporating visual context into audio understanding, particularly for native multimodal settings. Models with joint image–audio capability frequently outperform their audio-only or cascaded counterparts on speech, environmental sound, and music understanding tasks. For instance, Qwen- 2.5-omni exhibits a substantial performance gap between its audio+image configuration and audio- only variants (+15.55%, statistically significant across languages with paired t-test,p < 0.01), while the Gemma3n model, with slightly increased performance for Japanese, Malay, and Chinese. Sound category benefits the most from the visual in- formation (+11.7%for Qwen-2.5-omni and4.31% for Gemma3n among languages), which is ex- pected as the sound-related questions often require more contextual information to be correctly an- swered. These findings suggest that visual ground- ing can provide complementary semantic cues for acoustic reasoning, especially for ambiguous audio events where contextual scene information disam- biguates source identity, activity, or intent. Visual confusion occurs on cascaded modes. Despite the advantages of multimodal inputs, we identify a recurring failure mode in cascaded sys- tems that process modalities independently be- fore language reasoning. Cascaded variants of Gemma3n show degraded performance when im- age inputs are introduced alongside audio, indicat- ing that visual information can distract rather than assist the model’s reasoning process. This phe- nomenon is particularly visible in models adapted from speech-specific cases, where image condi- tioning sometimes lowers average accuracy across sound and music tasks. We hypothesize that fu- sion in cascaded pipelines induces representational competition: visual embeddings can be weakly cor- related with the target acoustic semantics. Such visual confusion highlights the limitation of late- fusion designs and motivates researchers with tighter cross-modal alignment and shared latent representations. Western languages perform better than east- ern languages.A clear geographic and linguistic disparity emerges across the multilingual bench- mark. Western languages, including German, En- glish, and Spanish, generally obtain higher average scores than East Asian languages such as Japanese and Chinese across most model families. The trend is particularly pronounced for sound and music understanding subtasks, where performance degra- dation in Japanese and Chinese remains substantial even among frontier proprietary models. This dis- crepancy likely reflects uneven multilingual audio pretraining distributions, differences in phonetic structure, and the relative scarcity of culturally di- verse non-speech acoustic supervision. Our find- ings suggest that multilingual multimodal compe- tence remains strongly biased toward high-resource Western language ecosystems. Model Input ModalityGermanEnglishSpanish ImageAudioSpeech / Sound / MusicAvg.Speech / Sound / MusicAvg.Speech / Sound / MusicAvg. Closed-source GPT-4o-mini-audio✘✔75.08 / 59.46 / 56.6363.7274.77 / 59.16 / 57.5363.8271.47 / 62.46 / 59.3464.42 GPT-4o-audio✘✔78.98 / 70.87 / 70.1873.5577.18 / 72.67 / 71.3973.7575.98 / 75.08 / 72.2974.45 GPT-4o † ✔✓ 65.77 / 69.07 / 57.2364.0268.77 / 70.57 / 59.6466.3367.57 / 68.17 / 59.0464.93 GPT-5-mini † ✘✓72.37 / 51.05 / 54.2259.2172.67 / 55.56 / 55.7261.3272.67 / 50.45 / 55.1259.41 GPT-5-2 † ✘✓71.17 / 48.05 / 51.5156.9172.67 / 51.95 / 50.0058.2169.07 / 54.05 / 51.5158.21 Open-source Phi-4-multimodal-instruct✘✔61.56/ 57.66 / 53.3157.5165.17 / 65.57 / 64.1665.6361.86/ 55.56 / 56.6358.01 Phi-4-multimodal-instruct † ✘✓45.95 / 61.86 / 59.9455.9250.75 / 65.47 / 63.2559.8249.55 / 64.56 / 56.9357.01 Phi-4-multimodal-instruct † ✔✓51.65 / 63.66 / 44.2853.2054.95 / 62.76 / 50.9056.2150.75 / 59.76 / 44.8851.80 Qwen-2.5-omni † ✘✓ 51.65 / 65.17 / 58.4358.4258.86 / 68.17/ 58.7361.9253.45 / 63.36 / 56.3357.71 Qwen-2.5-omni † ✔✓54.65 / 62.46 / 58.7358.6258.56 / 67.87 / 59.0461.8253.15 / 60.06 / 55.4256.21 Qwen-2.5-omni✘✔41.44 / 51.35 / 34.9442.8539.34 / 48.95 / 30.4239.5740.54 / 48.35 / 31.3340.07 Qwen-2.5-omni✔51.05 / 61.26 / 57.8356.7155.86 / 66.97 / 58.7360.5251.65 / 61.86 / 55.4256.31 Gemma3n † ✘✓ 56.76 / 69.67 / 52.4159.6160.06 / 66.67 / 56.9361.2258.86 / 68.47 / 53.0160.11 Gemma3n † ✔✓55.86 / 69.07/ 53.6159.5158.86 / 70.57 / 55.4261.6256.16 / 65.17/ 55.1258.81 Gemma3n✔✘41.44 / 48.35 / 46.3945.3942.94 / 51.65 / 48.1947.6039.94 / 47.75 / 45.4844.39 Gemma3n✘✔ 41.74 / 45.05 / 47.2944.6947.75 / 43.84 / 51.8147.8047.45 / 41.44 / 48.1945.69 Gemma3n✔39.94 / 48.05 / 46.3944.7943.24 / 48.35 / 48.4946.7040.24 / 48.95 / 47.8945.69 Gemma3n-EXAM 2 (Ours)✔66.37 / 62.76 / 59.9463.0264.56/ 62.16 / 59.9462.2264.26 / 58.86 / 61.1461.42 Model Input ModalityJapaneseMalayChinese ImAu Speech / Sound / MusicAvg.Speech / Sound / MusicAvg.Speech / Sound / MusicAvg. Closed-source GPT-4o-mini-audio✘✔74.47 / 62.16 / 56.0264.2271.17 / 60.06 / 56.3362.5275.08 / 62.76 / 56.9364.92 GPT-4o-audio✘✔79.28 / 74.17 / 68.0773.8478.68 / 73.27 / 69.8873.9477.48 / 72.97 / 70.4873.64 GPT-4o † ✔✓66.97 / 68.47 / 55.4263.6263.36 / 69.07 / 55.7262.7267.57 / 69.67 / 55.7264.32 GPT-5-mini † ✘✓ 71.77 / 48.65 / 53.9258.1169.97 / 51.65 / 54.8258.8171.47 / 50.15 / 51.9657.86 GPT-5-2 † ✘✓72.37 / 48.05 / 48.1956.2070.57 / 52.25 / 50.9057.9172.97 / 47.15 / 48.4956.20 Open-source Phi-4-multimodal-instruct✘✔59.16 / 51.05 / 51.8154.0151.05 / 44.14 / 49.7048.3060.66/ 52.85 / 54.5256.01 Phi-4-multimodal-instruct † ✘✓48.95 / 59.46 / 54.2254.2142.04 / 54.05 / 49.0148.4045.05 / 59.46 / 53.9252.81 Phi-4-multimodal-instruct † ✔✓ 50.15 / 58.56 / 42.1750.2946.25 / 57.06 / 39.1647.4954.05 / 57.66 / 43.0751.59 Qwen-2.5-omni † ✘✓51.65 / 63.36 / 53.0156.0148.65 / 59.16 / 54.8254.2157.36 / 65.77 / 57.8360.32 Qwen-2.5-omni † ✔✓48.95 / 58.56 / 50.0052.5048.05 / 56.46 / 53.9252.8156.16 / 61.26 / 55.1257.51 Qwen-2.5-omni✘✔ 38.74 / 48.35 / 33.4340.1738.14 / 52.55 / 32.5341.0739.64 / 49.85 / 32.8340.77 Qwen-2.5-omni✔ 50.15 / 59.16 / 48.1952.5046.55 / 56.46 / 52.7151.9057.96 / 63.96 / 56.9359.62 Gemma3n † ✘✓59.76 / 69.07 / 53.9260.9154.65/ 64.56 / 56.3358.5160.36 / 64.26/ 52.7159.11 Gemma3n † ✔✓60.06/ 65.77/ 49.7058.5153.45 / 63.66/ 55.1257.4158.56 / 63.96 / 51.8158.11 Gemma3n✔✘43.24 / 46.85 / 47.2945.7941.44 / 47.45 / 47.5945.4943.84 / 50.15 / 44.8846.29 Gemma3n✘✔ 41.74 / 42.34 / 45.7843.2934.53 / 43.24 / 49.1042.2943.84 / 42.34 / 47.8944.69 Gemma3n✔ 43.24 / 45.05 / 46.0844.7942.04 / 45.95 / 47.8945.2943.24 / 47.75 / 45.1845.39 Gemma3n-EXAM 2 (Ours)✔63.36 / 56.76 / 59.6459.9263.06 / 57.96 / 62.6561.2264.26 / 57.06 / 57.5359.62 Table 4: Multilingual and multimodal evaluation results (accuracy in %) on the EXAM 2 test set. Bold represents the best performance, whileunderlineindicates the second best performance for each category and each language within open-source models. † shows cascaded input of transcript and/or caption. We also mark✔and✓to differentiate the original audio waveform and textual input, respectively.✘ marks the unexploited modality. Cross-Linguistic Evaluation We also observe a notable domain similarity structure across lan- guages. Speech and sound performance display strong positive alignment (Pearson correlation≈ 0.80 ), implying that languages benefiting from ro- bust spoken-audio understanding also tend to per- form well on environmental sound tasks. Interest- ingly, Malay shows the strongest music understand- ing score (62.65%) among all languages, suggest- ing that music-related reasoning are different from others. Overall, the results reveal that multilingual audio understanding is not uniformly distributed across domains. Speech generalization transfers relatively well across languages, whereas sound and music understanding expose residual multilin- gual gaps, particularly for Eastern languages. 6 Ablation Study To assess the contribution of audio and vision modality to the overall performance of our model on the EXAM 2 benchmark, we conduct an abla- tion study by systematically removing either the audio or image input from our Gemma3n-EXAM 2 model on the EXAM 2 -test set across all six lan- guages in Section 6.1. Additionally, in Section 6.2, we compare the performance of our model trained on English-only data and all six languages train- ing data to investigate the impact of multilingual training on the performance of our model. Test Lang.Train Lang.Modality Accuracy (%) Speech / Sound / MusicAvg. EN ALLAudio+Image64.6 / 62.2 / 59.962.2 ENAudio+Image61.9 / 58.0 / 56.658.8 ALLAudio67.0 / 57.7 / 61.161.9 ALLImage40.5 / 43.5 / 43.442.5 DE ALLAudio+Image66.4 / 62.8 / 59.963.0 ENAudio+Image48.9 / 49.2 / 45.847.8 ALLAudio64.6 / 57.1 / 57.259.6 ALLImage38.7 / 42.9 / 41.641.1 ES ALLAudio+Image64.3 / 58.9 / 61.161.4 ENAudio+Image45.8 / 50.2 / 48.648.2 ALLAudio68.2 / 57.1 / 57.560.9 ALLImage40.5 / 45.3 / 41.942.6 JA ALLAudio+Image63.4 / 56.8 / 59.659.9 ENAudio+Image49.5 / 43.8 / 43.145.5 ALLAudio64.9 / 54.7 / 59.359.6 ALLImage38.4 / 40.5 / 40.740.0 MS ALLAudio+Image63.1 / 58.0 / 62.761.2 ENAudio+Image51.1 / 47.1 / 42.546.9 ALLAudio65.8 / 55.9 / 58.159.9 ALLImage41.7 / 45.0 / 40.142.3 ZH ALLAudio+Image64.3 / 57.1 / 57.559.6 ENAudio+Image51.4 / 45.6 / 40.445.8 ALLAudio67.0 / 53.8 / 54.558.4 ALLImage38.1 / 40.2 / 38.338.9 Table 5: Ablation study for our Gemma3n-EXAM 2 on the EXAM 2 -test by omitting different modalities on English-only and all six languages training data. 6.1 Ablation on Modality Overall, multimodal conditioning consistently provides the strongest performance. Averaged across all languages, multimodal module achieves 61.22% , outperforming audio-only (60.05%) and substantially exceeding image-only (41.23%). The improvement is particularly pronounced for envi- ronmental sound understanding, where multimodal inputs yield systematic gains over audio-only in- ference in every language, indicating that visual grounding supplies complementary contextual cues for disambiguating non-speech acoustic events. Interestingly, the relative benefit of modality differs across domains. Speech understanding achieves the best performance under audio-only setting, showing higher scores than audio+image in most languages. In contrast, multimodal fusion model provides clearer advantages for sound and, to a lesser extent, music understanding, where con- textual scene information can help infer semantic intent, source identity, or acoustic structure. The non-trivial performance of image-only models sug- gests that visual priors encode weak but exploitable correlations between scenes and expected acous- tic events. These findings reveal a complementary modality relationship: audio provides the primary semantic signal, while visual context improves ro- bustness and cross-domain generalization, particu- larly for non-speech audio understanding. 6.2 Ablation on Language We further examine the effect of training lan- guage composition by comparing models trained on English-only (EN) data against models trained on all six languages (ALL) under identical multi- modal configurations. Multilingual training pro- duces substantial and highly consistent improve- ments across all evaluation languages. Monolin- gual training leads to degradation−12.4points (48.83%averaged across languages). This pattern is universal across all six languages. Evidently, multilingual training benefits not only non-English languages but also English itself with a slight improvement of+3.4points, suggesting mul- tilingual exposure acts as a regularizer rather than introducing harmful language interference, improv- ing representation learning even for the dominant training language. A domain-wise breakdown fur- ther reveals that multilingual learning particularly strengthens sound and music understanding. For instance, in German evaluation, multilingual train- ing improves sound and music accuracies (+13.6 and+14.1gains respectively). Similar gains ap- pear across Japanese, Malay, and Chinese. We explain that multilingual supervision enriches the diversity of acoustic-language alignment during training, leading to representations that generalize better across culturally and linguistically heteroge- neous auditory phenomena. 7 Conclusion To summarize, we introduce EXAM 2 , a compre- hensive benchmark with in total135, 684multilin- gual and22, 614multimodal instances for evalu- ating audio understanding capabilities ofLALMs, LLMs,MLLMs, and cascaded systems. We con- duct extensive evaluations of open- and closed- source models across three audio domains (speech, sound, and music) and six languages (DE, EN, ES, JA, MS, and ZH). Results reveal that multi- modal conditioning significantly enhances audio understanding, particularly for non-speech sound tasks, while multilingual training substantially im- proves performance across all languages, especially for Eastern languages. Our proposed Gemma3n- EXAM 2 model, trained on our multilingual and multimodal training data, achieves competitive per- formance across all languages and audio categories, demonstrating the high quality of our benchmark for improving multilingual and multimodal audio understanding. Future work can explore more ef- ficient training methods and a wider range of lan- guages to further enhance the multilingual and mul- timodal capabilities for the EXAM 2 benchmark. Limitations There are several limitations to our current evalu- ation of the EXAM 2 benchmarks. In these bench- marks, we have prioritized accuracy over efficiency. Future evaluations should incorporate inference speed and deployment constraints to enable a more comprehensive assessment of model performance in real-world applications. Second, due to limited computational resources, we have only curated a subset of the AudioMCQ dataset to generate visual representations for training our model. Future work can explore the full dataset and consider more effi- cient training methods to leverage the entire dataset effectively. Additionally, our language coverage is limited to european and asian languages, a wider range of languages such as african language fami- lies can be included to better assess the multilingual capabilities of models on the EXAM 2 benchmarks. Ethic Statement The authors are not purposely creating or using any visual data that contains personally identifiable in- formation or offensive content. Generated images are solely used for training and evaluation purposes on the EXAM 2 benchmarks, and we have taken care to ensure that they do not contain any harmful or inappropriate content. Use of AI Tools We have used GPT-5-mini for translating the orig- inal English questions and choices into other five languages, and we have used GPT-image-2 for gen- erating the visual representations of the audio. The author also acknowledge the use of ChatGPT for assistance with grammar, punctuation, and vocabu- lary refinement, as well as debugging tasks. References Kabir Ahuja, Harshita Diddee, Rishav Hada, Milli- cent Ochieng, Krithika Ramesh, Prachi Jain, Ak- shay Nambi, Tanuja Ganu, Sameer Segal, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2023. MEGA: Multilingual evaluation of generative AI. In Proceedings of the 2023 Conference on Empir- ical Methods in Natural Language Processing, pages 4232–4267, Singapore. Inclusion AI, Biao Gong, Cheng Zou, Chuanyang Zheng, Chunluan Zhou, Canxiang Yan, Chunxiang Jin, Chunjie Shen, Dandan Zheng, Fudong Wang, and 1 others. 2025. Ming-omni: A unified multimodal model for perception and generation. arXiv preprint arXiv:2506.09344. Akari Asai, Sneha Kudugunta, Xinyan Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, and Hannaneh Hajishirzi. 2024. BUFFET: Benchmarking large language models for few-shot cross-lingual transfer. In Proceedings of the 2024 Conference of the North American Chap- ter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Pa- pers), pages 1771–1800, Mexico City, Mexico. Chen Chen, ZeYang Hu, Fengjiao Chen, Liya Ma, Jiax- ing Liu, Xiaoyu Li, Ziwen Wang, Xuezhi Cao, and Xunliang Cai. 2025. Uno-bench: A unified bench- mark for exploring the compositional law between uni-modal and omni-modal in omni models. arXiv preprint arXiv:2510.18915. Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, and 1 others. 2024. Qwen2-audio technical report. arXiv preprint arXiv:2407.10759. Soham Deshmukh, Benjamin Elizalde, Rita Singh, and Huaming Wang. 2023. Pengi: An audio language model for audio tasks. Advances in Neural Informa- tion Processing Systems, 36:18090–18108. Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen. 2020. Clotho: An audio captioning dataset. In ICASSP 2020-2020 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP), pages 736–740. IEEE. Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang. 2023. CLAP learning audio concepts from natural language supervision. In IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP 2023, Rhodes Island, Greece, June 4-10, 2023, pages 1–5. IEEE. Sreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Ku- mar, Zhifeng Kong, Sang-gil Lee, Chao-Han Yang, Ramani Duraiswami, Dinesh Manocha, Rafael Valle, and 1 others. 2026. Audio flamingo 3: Advancing au- dio intelligence with fully open large audio language models. Advances in Neural Information Processing Systems, 38:41819–41886. Sreyan Ghosh, Zhifeng Kong, Sonal Kumar, S Sak- shi, Jaehyeon Kim, Wei Ping, Rafael Valle, Di- nesh Manocha, and Bryan Catanzaro. 2025. Audio flamingo 2: An audio-language model with long- audio understanding and expert reasoning abilities. arXiv preprint arXiv:2503.03983. Sreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi, Chandra Kiran Evuru, S Sakshi, Oriol Ni- eto, Ramani Duraiswami, Dinesh Manocha, and 1 others. 2024. Compa: Addressing the gap in com- positional reasoning in audio-language models. In International Conference on Learning Representa- tions, volume 2024, pages 15486–15511. Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel. 2022. Audioclip: Extending clip to image, text and audio. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022, Virtual and Singapore, 23-27 May 2022, pages 976–980. IEEE. Haolin He, Xingjian Du, Renhe Sun, Zheqi Dai, Yujia Xiao, Mingru Yang, Jiayi Zhou, Xiquan Li, Zhengxi Liu, Zining Liang, and 1 others. 2025. Measuring audio’s impact on correctness: Audio-contribution- aware post-training of large audio language models. arXiv preprint arXiv:2509.21060. Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Liang Wang, Weizhu Chen, and 1 others. 2022. Lora: Low-rank adaptation of large language models. Iclr, 1(2):3. Chien-yu Huang, Ke-Han Lu, Shih-Heng Wang, Chi- Yuan Hsiao, Chun-Yi Kuan, Haibin Wu, Siddhant Arora, Kai-Wei Chang, Jiatong Shi, Yifan Peng, and 1 others. 2024. Dynamic-superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 12136–12140. IEEE. Laya Iyer, Angelina Wang, and Sanmi Koyejo. 2026. SCENEBench: An audio understanding benchmark grounded in assistive and industrial use cases. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Lin- guistics (Volume 1: Long Papers), pages 7123–7137, Rabat, Morocco. Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle, and Bryan Catanzaro. 2024. Audio flamingo: A novel audio language model with few- shot learning and dialogue abilities. arXiv preprint arXiv:2402.01831. Sonal Kumar, Simon Sedlácek, Vaibhavi Lokegaonkar, Fernando López, Wenyi Yu, Nishit Anand, Hyeong- gon Ryu, Lichang Chen, Maxim Plicka, Miroslav Hlavácek, William Fineas Ellingwood, Sathvik Udupa, Siyuan Hou, Allison Ferner, Sara Barahona, Cecilia Bolaños, Satish Rahi, Laura Herrera-Alarcón, Satvik Dixit, and 15 others. 2026. Mmau-pro: A chal- lenging and comprehensive benchmark for holistic evaluation of audio general intelligence. In Fortieth AAAI Conference on Artificial Intelligence, Thirty- Eighth Conference on Innovative Applications of Ar- tificial Intelligence, Sixteenth Symposium on Educa- tional Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, pages 22688–22697. AAAI Press. Yadong Li, Jun Liu, Tao Zhang, Song Chen, Tian- peng Li, Zehuan Li, Lijun Liu, Lingfeng Ming, Guosheng Dong, Da Pan, and 1 others. 2025. Baichuan-omni-1.5 technical report. arXiv preprint arXiv:2501.15368. Samuel Lipping, Parthasaarathy Sudarsanam, Konstanti- nos Drossos, and Tuomas Virtanen. 2022. Clotho- aqa: A crowdsourced dataset for audio question an- swering. In 2022 30th European Signal Processing Conference (EUSIPCO), pages 1140–1144. IEEE. Ziyang Ma, Yinghao Ma, Yanqiao Zhu, Chen Yang, Yi-Wen Chao, Ruiyang Xu, Wenxi Chen, Yuanzhe Chen, Zhuo Chen, Jian Cong, Kai Li, Keliang Li, Siyou Li, Xinfeng Li, Xiquan Li, Zheng Lian, Yuzhe Liang, Minghao Liu, Zhikang Niu, and 15 others. 2025. MMAR: A challenging benchmark for deep reasoning in speech, audio, music, and their mix. CoRR, abs/2505.13032. Jan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, and Soujanya Poria. 2024. Mustango: Toward controllable text- to-music generation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 8293–8316, Mexico City, Mexico. Zi Haur Pang, Xiaoxue Gao, Tatsuya Kawahara, and Nancy F Chen. 2026. Erm-minmaxgap: Bench- marking and mitigating gender bias in multilingual multimodal speech-llm emotion recognition. arXiv preprint arXiv:2603.21050. Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022.High- resolution image synthesis with latent diffusion mod- els. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695. S. Sakshi, Utkarsh Tyagi, Sonal Kumar, Ashish Seth, Ramaneswaran Selvakumar, Oriol Nieto, Ramani Du- raiswami, Sreyan Ghosh, and Dinesh Manocha. 2025. MMAU: A massive multi-task audio understanding and reasoning benchmark. In The Thirteenth Inter- national Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenRe- view.net. Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. 2024. Salmonn: Towards generic hearing abilities for large language models. In International Conference on Learning Representations, volume 2024, pages 16607–16629. Qwen Team. 2026. Qwen3. 5-omni technical report. arXiv preprint arXiv:2604.15804. Bin Wang, Xunlong Zou, Geyu Lin, Shuo Sun, Zhuohan Liu, Wenyu Zhang, Zhengyuan Liu, AiTi Aw, and Nancy F. Chen. 2025a. AudioBench: A universal benchmark for audio large language models. In Pro- ceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Compu- tational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 4297–4316, Albu- querque, New Mexico. Dingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang, Xueyuan Chen, Tianhua Zhang, and Helen Meng. 2025b. Mmsu: A massive multi-task spoken language understanding and reasoning benchmark. arXiv preprint arXiv:2506.04779. Shuting Wang, Jiejun Tan, Zhicheng Dou, and Ji-Rong Wen. 2025c. OmniEval: An omnidirectional and automatic RAG evaluation benchmark in financial domain. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 5726–5751, Suzhou, China. Benno Weck, Ilaria Manco, Emmanouil Benetos, Elio Quinton, George Fazekas, and Dmitry Bogdanov. 2024. Muchomusic: Evaluating music understand- ing in multimodal audio-language models. arXiv preprint arXiv:2408.01337. Valentin Werner. 2023. English and German pop song lyrics: Towards a contrastive textology. In Jour- nal for Language Technology and Computational Linguistics, Vol. 36 No. 1, pages 1–20, unknown. German Society for Computational Lingustics and Language Technology. Jin Xu, Zhifang Guo, Hangrui Hu, Yunfei Chu, Xiong Wang, Jinzheng He, Yuxuan Wang, Xian Shi, Ting He, Xinfa Zhu, and 1 others. 2025. Qwen3-omni technical report. arXiv preprint arXiv:2509.17765. Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu, Ziyue Jiang, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao, Chang Zhou, and Jingren Zhou. 2024. AIR-bench: Benchmarking large audio-language models via generative comprehension. In Proceed- ings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), pages 1979–1998, Bangkok, Thailand. Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, and 1 others. 2024. Mmmu: A massive multi-discipline multimodal un- derstanding and reasoning benchmark for expert agi. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition, pages 9556– 9567. Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, and Gra- ham Neubig. 2025. MMMU-pro: A more robust multi-discipline multimodal understanding bench- mark. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), pages 15134–15186, Vienna, Austria. Zihan Zhao, Yiyang Jiang, Heyang Liu, Yu Wang, and Yanfeng Wang. 2024. Librisqa: A novel dataset and framework for spoken question answering with large language models. IEEE Transactions on Artificial Intelligence. A Appendix A.1 Dataset Details A.1.1 Translation Guidelines When translating the original English questions and choices into other five languages, we first use the GPT-5-mini to follow the instruction in Figure 3. The translated results are then carefully reviewed by native speakers of the target languages (or re- searchers who are fluent in those languages for over 5 years) to ensure the accuracy, fluency, and appro- priateness of the translations. Any issues identified during the review process were addressed and cor- rected to ensure the highest quality of the translated questions and answer choices. Below is our specific guideline to ensure the quality and consistency of the translations: 1.Preserve Meaning: The translation should accurately convey the meaning of the original English text, including the question and all answer choices. 2.Maintain Format: The structure of the ques- tion and answer choices should be preserved in the translation, including the numbering and formatting. 3.Use Natural Language: The translation should use natural and fluent language that is appropriate for the target language, avoiding literal translations that may sound awkward or unnatural. 4.Cultural Relevance: If the original question contains culturally specific references or ex- amples, we may adapt them to be more rele- vant to the target language and culture while maintaining the original intent of the question. 5.Consistency Across Languages: We aimed to maintain consistency in the translation style and terminology across all languages to en- sure that the questions are comparable and that the evaluation is fair across different lin- guistic contexts. A.2 Model Overview A.2.1 Multilingual Examples Figure 7 and Figure 6 shows the multilingual trans- lation validated by native speakers on the EXAM 2 - test set. By not only checking the correctness and fluency of the translations, we also ensure that the Translate the following audio understanding/reasoning ques- tion, its choices, and its answer into Language. Keep technical/music notation unchanged. Keep the English words in the choices unchanged if they are asked which word appears first in the question or something similar. Make sure the translated answer matches one of the trans- lated choices. Question: question Choices: option 1. option 2. ... Answer: answer Figure 3: Translation prompt of GPT-5-mini. Lan- guage is replaced by the target language (German, Spanish, Japanese, Malay, and Chinese) when applying the prompt. A real-world and precise visual representation of the audio: option 1. , option 2. , .... Figure 4: Image generation prompt of GPT-image-2 and stable diffusion. translated answer in a consistent position with the original English answer. For instance, we manu- ally removed the redundant indexes of the given options, especially for the German and Spanish, as GPT-5-mini tends to automatically add the incon- sistent indexes in the translated choices, which may cause confusion for the model. Additionally, if the audio contains some specific words that are asked in the question, we keep those words unchanged in the translated questions and choices, as shown in Figure 6. This step aligns with our processed transcript and caption, as most of the audio con- tains English words, and we want to ensure the consistency between the content within languages. A.2.2 Visual Examples We provide visual examples of generated images for all three audio categories (speech, sound, and music) in Figure 8 and Figure 9. For images in Fig- ure 8a, Figure 8b, and Figure 8c, they are screen- shots manually created by authors, as the original images generated by stable diffusion/GPT-image-2 are not highly relevant to the audio content. How- ever, due to strong capability of visual understand- ing in environmental sounds, GPT-image-2 can generate more fine-grained images for sound un- derstanding questions, which is particularly helpful in our EXAM 2 -train, as shown in Figure 9. You are an expert in multimodal audio understanding. Let us solve a multiple-choice question together. Image 1 -> Choice 1 Image 2 -> Choice 2 Image 3 -> Choice 3 Image 4 -> Choice 4 Use ALL inputs: - Question - Choices - Transcript - Caption - Images Instructions: - Select the SINGLE best answer. - Return ONLY the exact choice text. - Do not output the choice number. - Do not explain. - Do not output anything else. Question: question Choices: choices_text Transcript: transcript Caption: caption Figure 5: Inference prompt for multimodal models in- cluding Phi4-multimodal-instruct, Qwen2.5-omni, and Gemma3n. A.3 Prompts Figure 3 shows the prompt we used for translat- ing the original English questions and choices into other five languages (DE, ES, JA, MS, and ZH) us- ing GPT-5-mini. For simplicity, prompt to generate our image using GPT-image-2 and stable diffusion is shown in Figure 4. When evaluating the performance of multimodal models, we use the same prompt as shown in Fig- ure 5 for all multimodal models. Specifc modali- ties (e.g., image-only or audio-only) are controlled by removing the corresponding input from the prompt. For instance, for image-only evaluation, we only keep the question, choices, and images in the prompt, while for audio-only evaluation, we only keep the question, choices, and original audio in the prompt. Cascaded evaluation is conducted by applying the transcript / caption instead of the audio input in the prompt. "audio_id": "./test-mini-audios/421fcbf8-7f60-4770-8923-512323e5efba.wav", "question": "Which word appears first", "choices": [ "push-ups", "pull-ups" ], "answer": "push-ups", "clotho": "a shuffled while snickers cat children talks", "transcript": "Push-ups, press-ups, pull-ups, sit-ups.", "answer_id": "1", "question_de": "Welches Wort erscheint zuerst?", "choices_de": [ "push-ups", "pull-ups" ], "answer_de": "push-ups", "question_es": "¿Qué palabra aparece primero?", "choices_es": [ "push-ups", "pull-ups" ], "answer_es": "push-ups", "question_ja": "どの単語が最初に現れますか?", "choices_ja": [ "push-ups", "pull-ups" ], "answer_ja": "push-ups", "question_ms": "Perkataan mana muncul terlebih dahulu", "choices_ms": [ "push-ups", "pull-ups" ], "answer_ms": "push-ups", "question_zh": "哪个单词先出现?", "choices_zh": [ "push-ups", "pull-ups" ], "answer_zh": "push-ups" Figure 6: Transcript and caption, multilingual transla- tion examples in json screenshot on our EXAM 2 . HyperparameterValue Data Typebfloat16 Learning Rate2e-4 Batch Size4 Gradient Accumulation Steps8 Training Epochs6 Lora Alpha32 Lora Dropout0.05 Table 6: Hyperparameters used for supervised fine- tuning of Gemma3n. A.4 Experiment Details We give additional detailed hyperparameters used in our experiments in Table 6. The backbone model for fine-tuning is Gemma3n-E4B, and we use the OmniLoRA method for efficient fine-tuning. "audio_id": "./test-mini-audios/69631267-f7ef-464e-8bc6-4f3e75e6fb6f.wav", "question": "Based on the given audio, which of the following best describes the sound environment?", "choices": [ "A forest ambience with natural bird sounds only", "A live orchestra performing in a theater", "An outdoor scene with electric guitar music and birds in motion", "A crowded marketplace with street vendors" ], "answer": "An outdoor scene with electric guitar music and birds in motion", "question_de": "Basierend auf der gegebenen Audioaufnahme, welche der folgenden Beschreibungen trifft die Klangumgebung am besten?", "choices_de": [ "Eine Waldatmosphäre mit ausschließlich natürlichen Vogelstimmen", "Ein Live-Orchester, das in einem Theater spielt", "Eine Außenszene mit E-Gitarrenmusik und Vögeln in Bewegung", "Ein belebter Marktplatz mit Straßenverkäufern" ], "answer_de": "Eine Außenszene mit E-Gitarrenmusik und Vögeln in Bewegung", "question_es": "Basándose en el audio proporcionado, ¿cuál de las siguientes opciones describe mejor el entorno sonoro?", "choices_es": [ "Un ambiente forestal con sonidos naturales de aves únicamente", "Una orquesta en vivo actuando en un teatro", "Una escena al aire libre con música de guitarra eléctrica y aves en movimiento", "Un mercado concurrido con vendedores ambulantes" ], "answer_es": "Una escena al aire libre con música de guitarra eléctrica y aves en movimiento", "question_ja": "与えられたオーディオに基づいて、次のうちどれが音の環境を最もよく表していますか?", "choices_ja": [ "自然の鳥の鳴き声のみが聞こえる森林の音環境", "劇場で演奏しているライブオーケストラ", "エレクトリックギターの音楽と鳥が飛び回る屋外のシーン", "屋台のある混雑した市場" ], "answer_ja": "エレクトリックギターの音楽と鳥が飛び回る屋外のシーン", "question_ms": "Berdasarkan audio yang diberikan, yang manakah berikut paling tepat menggambarkan persekitaran bunyi?", "choices_ms": [ "Suasana hutan dengan bunyi burung semula jadi sahaja", "Orkestra langsung yang tampil di teater", "Adegan luar dengan muzik gitar elektrik dan burung yang bergerak", "Pasar yang sesak dengan peniaga jalanan" ], "answer_ms": "Adegan luar dengan muzik gitar elektrik dan burung yang bergerak", "question_zh": "根据所提供的音频,下列哪一项最能描述该声音环境?", "choices_zh": [ "只有自然⻦鸣的森林氛围", "在剧院演出的现场管弦乐团", "伴有电吉他音乐和活动中⻦类声音的户外场景", "有街头小贩的拥挤市场" ], "answer_zh": "伴有电吉他音乐和活动中⻦类声音的户外场景" Figure 7: Multilingual translation examples in json screenshot on our EXAM 2 . (a) Visual examples of speech category on EXAM 2 evaluation dataset. The generated images are from choices prompts (from left to right): An hour; Two hours; Thirty minutes; Fifteen minutes. (b) Visual examples of sound category on EXAM 2 evaluation dataset. The generated images are from choices prompts (from left to right): Once; Twice; Three times; Four times. (c) Visual examples of music category on EXAM 2 evaluation dataset. The generated images are from choices prompts (from left to right): D#:min7/1; G#:min6(9,*1)/6; F#:maj7/1; C#:sus2(b7,*1)/b7. Figure 8: Visual examples manually created on EXAM 2 evaluation set. (a) Visual examples of speech category on EXAM 2 -Clotho training dataset. The generated images are from choices prompts (from left to right):A concert host introducing a band; A sports commentator describing a match; An auction announcer addressing a crowd; A train station departure announcer. (b) Visual examples of sound category on EXAM 2 -Clotho training dataset. The generated images are from choices prompts (from left to right):The rain and ocean waves; The hail and mountain stream; The wind and forest animals; The thunder and city traffic. (c) Visual examples of music category on EXAM 2 -Clotho training dataset. The generated images are from choices prompts (from left to right):The violin and the gong; The xylophone and the gong; The piano and the gong; The xylophone and the cymbals. Figure 9: Visual examples generated by GPT-image-2 on EXAM 2 -Clotho training set.