Paper deep dive
HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models
Feiyu Zhao, Yiming Chen, Wenhuan Lu, Daipeng Zhang, Xianghu Yue, Jianguo Wei
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/26/2026, 11:09:44 PM
Summary
HalluAudio is a large-scale, human-verified benchmark designed to evaluate hallucinations in Large Audio-Language Models (LALMs). It covers three major audio domains: speech, environmental sound, and music, featuring over 5,000 QA pairs. The benchmark utilizes adversarial prompts and mixed-audio conditions to diagnose failure modes such as acoustic grounding deficiencies, temporal reasoning errors, and semantic inconsistencies. The study evaluates 12 LALMs, revealing significant issues with yes/no bias, false refusals, and domain-specific reliability gaps.
Entities (9)
Relation Signals (6)
HalluAudio → coversdomain → Speech
confidence 100% · HalluAudio spans three major audio domains, speech, environmental sounds, and music
HalluAudio → coversdomain → Environmental Sound
confidence 100% · HalluAudio spans three major audio domains, speech, environmental sounds, and music
HalluAudio → coversdomain → Music
confidence 100% · HalluAudio spans three major audio domains, speech, environmental sounds, and music
HalluAudio → evaluates → LALM
confidence 100% · We therefore introduce HalluAudio, the first large-scale benchmark for evaluating hallucinations across speech, environmental sound, and music.
GPT-4o-Audio → isa → LALM
confidence 100% · including 2 proprietary models: GPT-4o-Audio
Qwen2-Audio → isa → LALM
confidence 100% · Qwen2-Audio (Chu et al., 2024)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks. However, hallucination, where models generate responses that are semantically incorrect or acoustically unsupported, remains largely underexplored in the audio domain. Existing hallucination benchmarks mainly focus on text or vision, while the few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth. We therefore introduce HalluAudio, the first large-scale benchmark for evaluating hallucinations across speech, environmental sound, and music. HalluAudio comprises over 5K human-verified QA pairs and spans diverse task types, including binary judgments, multi-choice reasoning, attribute verification, and open-ended QA. To systematically induce hallucinations, we design adversarial prompts and mixed-audio conditions. Beyond accuracy, our evaluation protocol measures hallucination rate, yes/no bias, error-type analysis, and refusal rate, enabling a fine-grained analysis of LALM failure modes. We benchmark a broad range of open-source and proprietary models, providing the first large-scale comparison across speech, sound, and music. Our results reveal significant deficiencies in acoustic grounding, temporal reasoning, and music attribute understanding, underscoring the need for reliable and robust LALMs.
Tags
Links
- Source: https://arxiv.org/abs/2604.19300v1
- Canonical: https://arxiv.org/abs/2604.19300v1
Trouble viewing inline? Open PDF directly →
Full Text
71,588 characters extracted from source content.
Expand or collapse full text
HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models Feiyu Zhao 1 * , Yiming Chen 2 * , Wenhuan Lu 1 , Daipeng Zhang 1 , Xianghu Yue 1† , Jianguo Wei 1 1 College of Intelligence and Computing, Tianjin University, China 2 ASUS Intelligent Cloud Services, Singapore 1 zhaofeiyu, wenhuan, zhangdaipeng, yuexianghu, jianguo@tju.edu.cn 2 MattYM_Chen@asus.com Abstract Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks. However, halluci- nation, where models generate responses that are semantically incorrect or acoustically un- supported, remains largely underexplored in the audio domain. Existing hallucination bench- marks mainly focus on text or vision, while the few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth. We therefore introduce HalluAudio, the first large-scale benchmark for evaluating hallucina- tions across speech, environmental sound, and music. HalluAudio comprises over 5K human- verified QA pairs and spans diverse task types, including binary judgments, multi-choice rea- soning, attribute verification, and open-ended QA. To systematically induce hallucinations, we design adversarial prompts and mixed-audio conditions. Beyond accuracy, our evaluation protocol measures hallucination rate, yes/no bias, error-type analysis, and refusal rate, en- abling a fine-grained analysis of LALM fail- ure modes. We benchmark a broad range of open-source and proprietary models, providing the first large-scale comparison across speech, sound, and music. Our results reveal significant deficiencies in acoustic grounding, temporal reasoning, and music attribute understanding, underscoring the need for reliable and robust LALMs. 1 1 Introduction Large Language Models (LLMs) have driven rapid progress in natural language processing, achieving strong performance across tasks such as reasoning, question answering, and multimodal understand- ing. Building on these advances, Large Audio- Language Models (LALMs) have emerged as a natural extension of LLMs to the audio domain. By * Equal contribution † Corresponding author 1 https://github.com/Feiyuzhao25/halluaudio leveraging large-scale corpora of speech, environ- mental sounds, and music, LALMs demonstrate impressive capabilities in speech recognition (Chu et al., 2023; Fang et al., 2025), sound question an- swering (Xu et al., 2025; Goel et al., 2025), and music understanding (Ghosh et al., 2025; Xiaomi, 2025). As LALMs are increasingly deployed in real-world applications, their accuracy and relia- bility have become critical. In particular, hallu- cinations remain a major concern, where models generate responses that are semantically incorrect or unsupported by the underlying audio. A substantial body of work has demonstrated that generative models often produce outputs that are fluent yet factually unsupported in both text- only (Li et al., 2023; Lin et al., 2022) and vi- sion–language settings (Guan et al., 2024; Wu et al., 2024). In contrast, hallucination in audio-centric models remains largely underexplored. Most ex- isting benchmarks focus on text or vision, while the few emerging studies in the audio domain are limited in scale, modality coverage, and task di- versity (Cheng et al., 2025). Current evaluations typically rely on small binary classification tasks and rarely probe critical failure modes such as re- sponse bias, refusal behavior, or multi-turn incon- sistency. As a result, the field of LALMs still lacks a dedicated large-scale benchmark for systemat- ically characterizing hallucination across speech, environmental sound, and music tasks. To bridge this gap, we introduce HalluAudio, the first large-scale human-verified benchmark suite specifically designed to evaluate hallucination in LALMs. HalluAudio spans three major audio do- mains, speech, environmental sounds, and music, and supports diverse task formats, including classi- fication, question answering, and open-ended gen- eration. To reliably elicit and measure hallucina- tions, we incorporate adversarial prompts, mixed audio conditions. To ensure the reliability of our evaluation suite, we also perform manual inspec- arXiv:2604.19300v1 [cs.SD] 21 Apr 2026 tion on the dataset. We curate HalluAudio through a five-step pipeline including data collection, task construction, adversarial augmentation, automated filtering, and multi-round human verification, re- sulting in a large-scale benchmark spanning three audio domains, dozens of task types, and over 5K carefully validated QA pairs. We evaluate a range of state-of-the-art open source and proprietary LALMs on HalluAudio. Comprehensive experiments show that hallucina- tion remains a systematic issue in current LALMs, even for tasks with a clear audio answer. We ob- serve consistent Yes/No biases, non-trivial false refusal behaviors, and domain-dependent failure patterns across speech, sound, and music. More importantly, strong performance on standard audio benchmarks does not necessarily imply robustness against hallucination, highlighting a gap between capability evaluation and reliability assessment. These findings underscore the need for targeted hallucination diagnostics and validate HalluAudio as an effective tool for analyzing fine-grained fail- ure modes beyond aggregate accuracy. Our main contributions are as follows: • A large-scale human-verified benchmark for audio hallucination. We introduce Hallu- Audio, the first large-scale human-verified au- dio hallucination benchmark spanning speech, environmental sounds, and music, with thou- sands of QA pairs per domain covering both audio understanding and audio-grounded gen- eration scenarios. •A diverse suite of hallucination-inducing task designs. HalluAudio includes binary and multi-class classification, audio QA, at- tribute verification, comparative reasoning, and open-ended generation. We construct ad- versarial prompts, mixed-audio inputs, and positives/negatives to systematically trigger and measure hallucinations. •A multi-dimensional empirical analysis of hallucination in LALMs. We evaluate a broad set of open-source and proprietary LALMs and report detailed hallucination pat- terns across tasks and modalities. Our analy- sis incorporates accuracy, hallucination rate, Yes/No bias, error-type breakdown, and re- fusal rate, revealing critical reliability gaps in current LALMs. 2 Related Works Large-Audio Language ModelsInspired by the success of LLMs, recent work integrates audio rep- resentations with LLMs, giving rise to LALMs that process speech, environmental sounds, and music to generate textual responses and reason- ing outputs. Early efforts such as AudioLM (Bor- sos et al., 2023) demonstrated that discretized audio tokens can be effectively modeled with language modeling techniques, enabling coher- ent continuation of speech and music. Subse- quent models, including SALMONN (Tang et al., 2024), Qwen2-Audio (Chu et al., 2024), and Audio Flamingo (Kong et al., 2024), further align audio encoders with LLMs to support tasks such as ASR, audio question answering, captioning, and music understanding. More recent studies extend LALMs to finer-grained music understanding (Huang et al., 2022) and improved instruction following (Frieske and Shi, 2024). Overall, LALMs are rapidly evolv- ing from transcription-focused systems to general audio reasoning models. Although several bench- marks have been proposed to evaluate various ca- pabilities of LALMs (Yang et al., 2024; Chen et al., 2024, 2026b,a), their reliability, particularly with respect to hallucination, remains underexplored. Hallucinations in Audio Tasks Hallucination has been extensively studied in text (Bang et al., 2025) and vision (Cao et al., 2024) domains, re- vealing systematic failures in grounding, object hallucination, and cross-modal consistency. In con- trast, hallucinations in the audio domain remain largely underexplored. Frieske and Shi (2024) pro- vides an early analysis of hallucination in ASR, showing that metrics such as WER fail to detect fluent but semantically irrelevant outputs and ex- posing vulnerabilities to misleading acoustic cues. AHa-Bench (Cheng et al., 2025) extends this line of inquiry to LALMs through a small-scale bi- nary QA benchmark. While informative, existing benchmarks are limited in dataset scale, task diver- sity, and diagnostic depth, leaving critical failure modes such as response bias and refusal behav- ior. The field still lacks a unified taxonomy, con- trolled contrastive audio pairs, multi-format eval- uation tasks, and large-scale human-annotated as- sessments. HalluAudio addresses these gaps by introducing the first large-scale, multi-domain, and multi-dimensional benchmark for hallucination in LALMs. Dataset ConstructionEvaluation SetupHallucination Analysis lZero-shot Inference lUnified Evaluation Protocol Speech Sound Music Clean Speech Overlapped Speech Speed/loudness Single Sound Co-occurring Sounds Absent Sound Multi-singer Genre Multi-instrument Speech–music Binary Yes/No Numeric Counting Multi-label Classify Intentionally Invalid queries Template-based, Parameterized Prompts Domain-aware Instantiation Controlled Answer Formats Prompt LALMs Two sounds are played. Which sound is louder? I cannot determine which sound is louder based on the information provided as I do not have access to the audio... The second sound is louder. Generated Response Expected Response Output Type Normalizer Validity Checker Behavior Analyzer Free-form Output Audio Task Ground Truth Validity Decision Incorrect Output Yes/No Bias False Refusal Other Errors 【Yes】【No】【Number】 【Refusal】【Other】 Figure 1: Overview of the HalluAudio framework. HalluAudio combines controlled multi-domain audio construction, unified prompting, and structured output validation to systematically analyze hallucination in LALMs. 3 HalluAudio Benchmark 3.1 Overview of HalluAudio The overall evaluation pipeline of HalluAudio is il- lustrated in Figure 1, which follows a modular, end- to-end process designed to systematically elicit, measure, and analyze hallucinations in LALMs. Specifically, given curated audio inputs from three domains: speech, sound, and music, we con- struct diverse task instances via template-based, parameterized prompts with domain-aware instan- tiation and controlled answer formats. These au- dio–prompt pairs are evaluated under a unified zero- shot protocol across a suite of LALMs, producing textual outputs in heterogeneous forms, including binary decisions, numeric counts, and free-form responses. Finally, model outputs are then nor- malized into structured types and validated against task definitions and ground-truth audio evidence through an automated evaluation engine. Based on the validated outcomes, we conduct fine-grained hallucination analysis across domains and tasks, covering Yes/No bias, false refusals, and other er- ror patterns. In this work, we define audio hallucination op- erationally as a model-generated claim that is not supported by the acoustic evidence in the input. This includes three representative cases: (1) fab- rication, where the model asserts the presence of non-existent audio events; (2) evidence contradic- tion, where the response conflicts with determinis- tic acoustic structure; and (3) unjustified affirmative bias, where the model produces positive responses despite insufficient or absent evidence. This defi- nition explicitly distinguishes hallucination from general capability errors, such as failures due to reasoning complexity or ambiguous inputs. Based on this definition, HalluAudio organizes evaluation into three hallucination-sensitive dimen- sions aligned with task design: (1) Structural and temporal hallucinations, evaluated by Temporal Comparison tasks, where models make incorrect claims about order, timing, or quantitative acoustic properties; (2) Perceptual hallucinations, captured by Recognition tasks, involving incorrect asser- tions about the presence or attributes of sounds or musical elements; and (3) Semantic hallucinations, assessed through Consistency tasks, where model responses become internally inconsistent or unsup- ported under controlled input perturbations. Our objective is to provide a systematic, multi-domain evaluation protocol for analyzing hallucination be- haviors in modern LALMs. DomainDataset SpeechCommon Voice (Ardila et al., 2020) SoundFSD50K (Fonseca et al., 2021) Music GTZAN (Sturm, 2012) Mridangam Strokes (Turian et al., 2022) Mridangam Tonics (Turian et al., 2022) Table 1: Audio sources used in HalluAudio. Which sound is louder, the first or the second? First/Second. Do car and wind appear together in the recording? Does this recording contain a thunder sound? Does“dsh”appear before“ca”in the recording? What does the male say in the recording? How many times does the word“ok”appear in the recording? Does“apple”appear before ”pink”in the recording? Does the audio contain the specified instrument? Yes/No. None. (A female is talking.) 1/2/3. Yes/No. Yes/No. (Stroke music sound.) Yes/No. Ok, ok, ok. Apple, pink. Yes/No. (A train passing by.) (Car driving in winds.) (Dsh.)(Ca.) Figure 2: Detailed task composition of the HalluAudio dataset across speech, sound, and music domains. 3.2 Dataset Construction We construct HalluAudio through a controlled, re- producible pipeline designed to elicit and diagnose hallucination behaviors in LALMs. Step 1. Audio selection We curate audio clips from speech, environmental sound, and music cor- pora with reliable annotations. Clips are selected to cover diverse acoustic conditions, including multi- speaker speech, event-rich sound, and structured musical segments. Audio sources are summarized in Table 1. Step 2. Template-based prompt generationFor each task, we design parameterized prompt tem- plates with slot variables. Slots are instantiated using clip annotations for valid queries or delib- erately mismatched attributes for invalid queries, producing large-scale audio–question pairs. Step 3. Contrastive and adversarial construc- tion We generate paired instances by minimally modifying prompts or audio attributes, ensuring controlled positive/negative contrasts that isolate hallucination triggers. Audio attributes cover tem- poral order, event presence, loudness, counting, and musical structure, while prompt modifications include adversarial negatives, invalid queries, and attribute-level perturbations. Step 4. Validation and quality controlWe first generate a large pool of candidate instances using automated scripts. Each instance then undergoes three rounds of human verification involving two independent annotators and one senior reviewer. Annotators are instructed to check: (1) alignment between audio content and the question, (2) correct- ness of the ground-truth answer, and (3) whether the instance satisfies the intended hallucination- triggering condition. Disagreements are resolved through majority voting and adjudication by the se- nior reviewer. QA pairs with persistent ambiguity or inconsistent interpretations are revised or dis- carded to ensure high dataset reliability. The anno- tators are three PhD-level researchers specializing in speech and audio processing, ensuring domain expertise during verification. Step 5. Packaging and balancing In final step, we further balance the dataset across domains, task types, and hallucination categories to avoid distri- butional bias and ensure fair evaluation of different hallucination behaviors. 3.3 Dataset Statistics As illustrated in Figure 2, HalluAudio covers three audio domains with a diverse set of task categories and balanced sample distributions. The speech do- main contains the largest variety of fine-grained tasks, including temporal reasoning, transcription Domain CategoryCount Prompt Template#O-QA Speech Overlap Check189 Do the two speakers’ voices overlap in the recording?× Word Order245Does the word “[A]” appear before the word “[B]” in the recording?× Binary Count156 Does the speaker say the word “[A]” [X] times in the recording? × Exact Count178 How many times does the word “[A]” appear in the recording?✓ Invalid Gender172 What does the (fe)male say in the recording?✓ Invalid Noise192 What does the speaker say in the recording?✓ Deletion Match303 Does the speech recording match the transcription: “[A]”?× Swap Match310 Does the speech recording match the transcription: “[A]”?× Speed Compare230 Which instance of “[A]” was spoken faster, the first or the second?✓ Loudness Compare225 Which instance of “[A]” was spoken louder, the first or the second?✓ Sound Overlap Check254 Do “[LABEL1]” and “[LABEL2]” overlap in the recording?× Sound Order300 Does “[LABEL1]” appear before “[LABEL2]” in the recording?× Sound Presence260 Does the recording contain a “[LABEL]” sound?× Sound Coexist300 Do “[LABEL1]” and “[LABEL2]” appear together in the recording? × Mismatched Query257 Does this recording contain a “[RANDOM_LABEL]” sound?× Multi-label Check287 Are there multiple sound types in this recording?× Loudness Compare300 Which sound is louder, the first or the second?✓ Music Genre Match291 Is this audio clip [LABEL] music?× Instrument Match258 Does the audio contain the specified instrument or stroke label?× Speech or Music128 Is this audio clip speech or music?✓ Instrument Identify233 Is the stroke or tonic type [LABEL]?× Loudness Compare252 Which sound is louder, the first or the second?✓ Music Order297 Does “[LABEL1]” appear before “[LABEL2]” in the recording?× Instrument Count300 How many sound types in this recording?✓ Total-5720 -- Table 2: Detailed task composition of the HalluAudio dataset across speech, environmental sound, and music domains. #O-QA: open-ended QA pairs. consistency checks, comparative judgments, and invalid or underspecified queries, reflecting the complexity of speech-based hallucination behav- iors. Environmental sound tasks emphasize sound presence, co-occurrence, and adversarial negative queries, while music tasks focus on instrument and genre identification, comparative reasoning, and cross-domain invalid prompts. HalluAudio explicitly incorporates contrastive and adversarial constructions to enable targeted hal- lucination diagnosis. Contrastive tasks account for 2,662 out of 5,720 QA pairs, while explicitly adver- sarial or invalid queries account for 621 QA pairs. In total, 57.4% of the dataset is designed to probe hallucination through controlled perturbations or absence of evidence, distinguishing HalluAudio from conventional audio QA datasets that primar- ily contain valid and answerable queries. Task Composition and Statistics Table 2 presents a detailed breakdown of the HalluAu- dio dataset. The benchmark spans three domains, speech, environmental sound, and music, and cov- ers a diverse set of task categories designed to probe temporal reasoning, counting, matching, compar- ison, and invalid or underspecified queries. Each task is instantiated using parameterized prompt templates with slot variables, enabling large-scale construction while maintaining precise alignment between audio evidence and ground-truth answers. Invalid query categories intentionally lack suffi- cient audio support and are used to diagnose hallu- cination behaviors such as overconfident guessing or inappropriate refusals. Benchmarks#Ds-D#H-E#O-QAScale USMQ (Kuan et al., 2024)× × × ∼30K Match (Kuan and Lee, 2025) × × × >15K Avhbench (Sung-Bin et al., 2025) × × × ∼5K AHa-Bench (Cheng et al., 2025) ×✓ × ∼1K HalluAudio (ours)✓ >5K Table 3: Comparison of different hallucination bench- marks for LALMs. #Ds-D: Domain-specific Design. #H-E: Number of manually verified QA pairs. #O-QA: Open-ended QA. Model Temporal ComparisonRecognitionConsistency Average overlaporderspeedloudnessexactbinarynoisegender match_smatch_d Qwen-Audio55.7846.8151.5048.9744.5843.571.860.1314.9413.4132.16 Qwen2-Audio50.0051.463.350.006.5747.310.0018.9758.7150.8328.72 Llama-Omni41.2638.110.000.043.9237.280.110.0030.9630.9618.26 Llama-Omni250.3855.3821.8720.0248.7256.3999.9312.9417.3626.3640.94 Kimi-Audio52.2751.7329.7846.4948.8758.538.4082.5715.1813.4740.73 Phi-4-Multimodal 54.0757.4019.4029.3569.1559.320.00 0.8539.7525.4235.47 Pengi50.0034.10 0.000.0020.5839.260.000.0011.7411.5116.72 MiMo-Audio96.3079.5924.7847.5657.8755.132.620.5891.2272.7352.84 Step-Audio-299.4761.2250.0050.2248.3151.921.563.4914.152.5338.29 GPT-4o-Audio57.6779.184.3513.3380.3465.386.043.5175.1270.2045.51 Gemini-2.5-Flash9.8438.591.740.002.2635.762.147.5656.5711.5116.60 Table 4: Classification accuracy (%) on HalluAudio in speech domain. match_s: swap match. match_d: deletion match. Comparison with Other Benchmarks.Table 3 summarizes representative hallucination bench- marks for LALMs. Existing benchmarks such as USMQ (Kuan et al., 2024), Match (Kuan and Lee, 2025), and Avhbench (Sung-Bin et al., 2025) are of moderate scale and lack domain-specific task de- sign, manual verification, or open-ended QA. AHa- Bench (Cheng et al., 2025) provides manually veri- fied QA pairs but is limited in scale and restricted to binary hallucination detection. In contrast, Hal- luAudio adopts domain-specific task formulations across speech, environmental sound, and music, covering temporal, perceptual, and structural rea- soning, and supports both binary and open-ended QA at a larger scale (>5K QA pairs), enabling more fine-grained analysis of hallucination behav- iors across domains. 4 Evaluation 4.1 Examined Models We evaluate the performance of 12 LALMs, in- cluding 2 proprietary models: GPT-4o-Audio (Hurst et al., 2024) and Gemini-2.5-Flash (Co- manici et al., 2025), as well as 10 representative open-source models: Qwen-Audio (Chu et al., 2023), Qwen2-Audio (Chu et al., 2024), Qwen2.5- Omni (Xu et al., 2025), Llama-Omni (Grattafiori et al., 2024),Llama-Omni2 (Fang et al., 2025), Kimi-Audio (Ding et al., 2025), Phi- 4-Multimodal (Abouelenin et al., 2025), Au- dio Flamingo 3 (Goel et al., 2025), Music- Flamingo (Ghosh et al., 2025), Pengi (Deshmukh et al., 2023), MiMo-Audio (Xiaomi, 2025), Step- Audio-2 (Wu et al., 2025). All models are evaluated with three independent runs, and reported results are averaged to ensure statistical stability. 4.2 Evaluation Metrics We evaluate hallucination behaviors in LALMs us- ing metrics beyond standard accuracy, capturing both response correctness and bias patterns. Accuracy. Accuracy measures the fraction of prompts with well-defined ground truth that are answered correctly by LALMs: Accuracy d = 1 |P d | X p∈P d 1ˆy p = y p ,(1) whereP d denotes prompts in domaind,y p is the ground-truth answer, andˆy p is the model prediction. This allows fine-grained analysis across domains and reasoning skills. Accuracy captures explicit hallucination cases where model predictions con- tradict clearly defined ground-truth audio evidence. Yes/No Bias Test.This test diagnoses systematic bias in binary responses using three complementary measures. Yes-p Ratio = P p∈P binary 1ˆy p = Yes |P binary | ,(2) Unrelated Ratio = P p∈P binary 1ˆy p ̸= y p ∧ ˆy p P p∈P binary 1ˆy p ̸= y p ,(3) Conditional Accuracy = P p∈P binary 1ˆy p = y p |P binary | ,(4) whereP binary is the set of Yes/No prompts, andp∈ P binary splits by predicted class. These measures Model Temporal ComparisonRecognitionConsistency Average overlaporderloudnessmulti_labelpresencecoexistmismatch Qwen2.5-Omni87.9788.8992.7775.3587.9761.6898.1784.69 Audio Flamingo-363.5398.7095.2513.2726.9572.1383.6764.79 Pengi23.9023.908.6715.4734.3420.5046.0724.69 Kimi-Audio28.9039.3045.5011.8093.6145.806.4038.76 Qwen2-Audio64.6394.4795.7848.2397.6059.4111.4367.36 MiMo-Audio94.8891.0093.0042.8671.5489.0095.3382.52 Step-Audio-272.4492.6757.6726.8399.2397.0062.6572.64 GPT-4o-Audio66.5385.5272.6756.1095.3581.0039.8471.00 Gemini-2.5-Flash25.5327.9640.7751.4510.12 41.3778.2439.35 Table 5: Classification accuracy (%) on HalluAudio in sound domain. Model Temporal ComparisonRecognitionConsistency Average ordercount_scount_t loudness instru_idgenres_or_m match_s match_t Qwen2.5-Omni49.3825.8014.8495.4043.9262.16100.0034.0733.0750.96 Audio Flamingo350.0049.1748.53100.0070.5550.0050.0044.0760.5758.10 Pengi0.0051.3729.230.0032.5563.500.0028.8349.9028.38 Kimi-Audio100.0029.5320.17100.0032.9050.450.0033.7741.6045.38 Qwen2-Audio0.004.205.230.0053.3051.8550.0051.0732.3727.56 Music-Flamingo50.0028.1120.42100.0053.8050.00100.0030.8073.6056.30 MiMo-Audio100.0069.3367.33100.0067.8177.66100.0051.9648.0475.79 Step-Audio-2100.0038.0036.670.00 77.6878.6992.9748.6050.2858.10 GPT-4o-Audio99.6621.3314.6775.4067.8179.8699.2228.4936.8758.15 Gemini-2.5-Flash13.4512.7011.028.6420.9145.8787.5040.6742.8631.51 Table 6: Classification accuracy (%) on HalluAudio in music domain. count_s: instrument count in stroke. count_t: instrument count in tonic. instru_id: instrument identify. s_or_m: speech or music. match_s: instrument match in stroke. match_t: instrument match in tonic. diagnose systematic affirmative bias, where models tend to produce unsupported positive responses despite insufficient or contradictory evidence, a common form of hallucination behavior. False Refusal Rate (FRR). FRR captures cases where LALMs abstain despite a valid prompt: FRR d = |p∈ P d : ˆy p ∈R| |P d | ,(5) whereRdenotes refusal responses. FRR reflects over-conservative hallucination behavior, where models fail to respond despite the presence of suffi- cient evidence, indicating breakdowns in evidence- based decision making. 4.3 Results Accuracy Analysis. Tables 4, 5, and 6 report classification accuracy on HalluAudio across the speech, environmental sound, and music domains. Bold and underlined values indicate the maximum and minimum. Overall, hallucination behavior is highly task- and domain-dependent, and no single model exhibits uniformly robust performance. In the speech domain, structural hallucinations strongly impact temporal tasks such as counting and ordering. MiMo-Audio and Step-Audio-2 perform well, while Qwen-Audio, Llama-Omni, and Pengi remain below 50% on several subtasks. Semantic hallucinations appear in transcription- related prompts: Phi-4-Multimodal degrades under noise and gender perturbations, and Kimi-Audio shows inconsistent recognition. GPT-4o-Audio is more balanced overall but still struggles with fine- grained perceptual judgments. In the sound domain, perceptual and structural hallucinations are amplified in multi-label, coex- istence, and adversarial-negative settings. MiMo- Audio and Qwen2.5-Omni consistently outperform Figure 3: Yes/No Bias analysis of speech and environmental sound. From left to right: Yes-pred Ratio, Unrelated Error Ratio, and Conditional Accuracy. Higher Yes-pred Ratios combined with low conditional accuracy indicate strong affirmative bias rather than evidence-grounded binary reasoning. others on temporal comparison tasks, while Au- dio Flamingo3 and Pengi struggle with multi-label sound recognition. Loudness comparison remains unstable across models, and random-false prompts expose divergent behaviors: some models confi- dently hallucinate sound presence, while others over-refuse despite clear acoustic evidence. In the music domain, models frequently exhibit semantic hallucinations in genre and instrument identification under ambiguous audio. GPT-4o- Audio and MiMo-Audio perform relatively well, while Pengi and Qwen2-Audio approach near- random accuracy. Structural and temporal errors persist in stroke and tonic counting, with MiMo- Audio leading and Qwen2-Audio and Gemini-2.5- Flash falling below 15%. Perceptual hallucinations also occur in single- vs. multi-source detection, where Qwen2.5-Omni and GPT-4o-Audio near per- fect accuracy, unlike several open-source models. Across the three domains, both open-source and proprietary LALMs show distinct, non-uniform hal- lucination patterns, highlighting the need for fine- grained cross-domain evaluation beyond aggregate accuracy. Yes/No Bias Analysis. Figure 3 examines Yes/No bias from three complementary views. The Yes prediction ratio shows that Qwen2.5-Omni, Figure 4: False Refusal Rate in speech and environmen- tal sound domains. Each heatmap reports task-specific refusal frequencies for different models, where higher values indicate a stronger tendency to refuse answering despite the presence of a valid ground-truth answer. Qwen2-Audio, and Kimi-Audio consistently over- predict affirmative answers across speech and envi- ronmental sound, with the strongest skew on order- ing and counting tasks, indicating a largely domain- agnostic bias. The unrelated error ratio suggests that affirma- tive bias does not necessarily translate into seman- tically unrelated errors: Qwen-series models main- tain relatively low unrelated ratios, while Pengi and Audio Flamingo3 deteriorate sharply under swap and deletion perturbations, reflecting weak- ened grounding. Finally, conditional accuracy reveals asymmet- ric decision behavior—models with strong affir- mative bias achieve higher accuracy on positive cases but underperform on negatives, whereas Phi- 4-Multimodal exhibits a more balanced yet con- servative profile. Overall, affirmative bias is most pronounced in environmental sound tasks, where negative evidence must be inferred from the ab- sence of acoustic events. Detailed analysis of mu- sic domain can be found in the appendix B. False Refusal Analysis. Figure 4 reports false refusal behaviors in speech and environmental au- dio domains. The behavior of music domain can be found in the appendix B. In the speech domain, false refusals concentrate on structurally demand- ing tasks such as counting, speed, and loudness comparison. Qwen2-Audio shows extreme refusal on count and speed, while Gemini-2.5-Flash ex- hibits consistently high refusal across most tasks, indicating over-conservative abstention. In contrast, Phi-4-Multimodal, Kimi-Audio, MiMo-Audio, and Step-Audio-2 maintain near-zero refusal, suggest- ing stronger grounding under clear acoustic evi- dence. In environmental sound domain, refusal behavior reflects perceptual uncertainty rather than task com- plexity. Qwen2-Audio spikes on coexist, and loud- ness, and Gemini-2.5-Flash again refuses broadly, whereas Qwen2.5-Omni, MiMo-Audio, and Step- Audio-2 remain stable even under adversarial- negative settings. Overall, false refusals are not uniform safety be- haviors but a distinct hallucination mode, arising from failures in structural reasoning, perceptual confidence, or overly conservative decision poli- cies, with Gemini-2.5-Flash and Qwen2-Audio ex- hibiting the most severe over-refusal tendencies. Overall, false refusals form a distinct halluci- nation mode, driven by reasoning, perceptual, or conservative policy failures, with Gemini-2.5-Flash and Qwen2-Audio showing the most severe cases. 4.4 Qualitative Comparison Across domains, models exhibit distinct and of- ten inconsistent hallucination profiles, indicat- ing strong domain- and structure-dependence in LALMs. In the speech domain, performance is stable on basic recognition but degrades sharply on structurally demanding tasks such as counting, ordering, and speed comparison. Qwen2-Audio and Gemini-2.5-Flash frequently default to refusal, whereas Phi-4-Multimodal and MiMo-Audio main- tain more robust structural grounding. In contrast, Kimi-Audio tends to over-assert on binary and in- valid queries, producing confident yet weakly sup- ported responses. In the environmental sound domain, percep- tual hallucinations dominate. Qwen2.5-Omni and MiMo-Audio show more balanced behavior on event presence and co-occurrence, while Audio Flamingo3 and Pengi degrade under multi-label and adversarial-negative settings. Several models exhibit affirmative bias without sufficient ground- ing, highlighting weakness in absence-based rea- soning. The music domain shows the greatest model di- vergence. GPT-4o-Audio and MiMo-Audio are more robust on high-level semantic queries, while most models struggle with structural and tempo- ral reasoning such as counting and order. Qwen2- Audio and Gemini-2.5-Flash exhibit near-random performance or abrupt refusal spikes, revealing fragile musical structure understanding. Overall, no model is consistently robust across domains and hallucination types, highlighting the need for multi- domain, multi-metric evaluation of LALMs. 5 Conclusion This paper presents HalluAudio, a diagnostic benchmark for systematically evaluating halluci- nation behaviors in LALMs across speech, envi- ronmental sound, and music domains. Through fine-grained accuracy analysis, Yes/No bias testing, and FRR evaluation, we demonstrate that halluci- nations in audio understanding manifest in diverse and previously underexplored forms beyond incor- rect predictions. Our findings highlight the limi- tations of accuracy-centric evaluation and under- score the need for reliability-oriented benchmarks to guide the development of more robust LALMs. References Abdelrahman Abouelenin, Atabak Ashfaq, Adam Atkin- son, Hany Awadalla, Nguyen Bach, Jianmin Bao, Alon Benhaim, Martin Cai, Vishrav Chaudhary, Con- gcong Chen, and 1 others. 2025. Phi-4-mini tech- nical report: Compact yet powerful multimodal lan- guage models via mixture-of-loras. arXiv preprint arXiv:2503.01743. Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gre- gor Weber. 2020. Common voice: A massively- multilingual speech corpus. In Proceedings of the twelfth language resources and evaluation confer- ence, pages 4218–4222. Yejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn, Tara Fowler, Cheng Zhang, Nicola Can- cedda, and Pascale Fung. 2025. Hallulens: Llm hal- lucination benchmark. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 24128– 24156. Zalán Borsos, Raphaël Marinier, Damien Vincent, Eu- gene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, and 1 others. 2023. Audiolm: a language modeling approach to audio generation. IEEE/ACM transactions on audio, speech, and lan- guage processing, 31:2523–2533. Qingxing Cao, Junhao Cheng, Xiaodan Liang, and Liang Lin. 2024. Visdiahalbench: A visual dia- logue benchmark for diagnosing hallucination in large vision-language models. In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 12161–12176. Jianan Chen, Xiaoxue Gao, Tatsuya Kawahara, and Nancy F Chen. 2026a. Loasr-bench: Evaluating large speech language models on low-resource automatic speech recognition across language families. arXiv preprint arXiv:2603.20042. Yiming Chen, Xianghu Yue, Xiaoxue Gao, Chen Zhang, Luis Fernando D’Haro, Robby T. Tan, and Haizhou Li. 2024. Beyond single-audio: Advancing multi- audio processing in audio large language models. In Findings of the Association for Computational Lin- guistics: EMNLP 2024, pages 10917–10930, Miami, Florida, USA. Association for Computational Lin- guistics. Yiming Chen, Xianghu Yue, Chen Zhang, Xiaoxue Gao, Robby T Tan, and Haizhou Li. 2026b. Voicebench: Benchmarking llm-based voice assistants. Transac- tions of the Association for Computational Linguis- tics, 14:378–398. Xize Cheng, Dongjie Fu, Chenyuhao Wen, Shannon Yu, Zehan Wang, Shengpeng Ji, Siddhant Arora, Tao Jin, Shinji Watanabe, and Zhou Zhao. 2025. AHa- bench: Benchmarking audio hallucinations in large audio-language models. In The Thirty-ninth Annual Conference on Neural Information Processing Sys- tems Datasets and Benchmarks Track. Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, and 1 others. 2024. Qwen2-audio technical report. arXiv preprint arXiv:2407.10759. Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shil- iang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. 2023. Qwen-audio: Advancing universal audio understanding via unified large-scale audio- language models. arXiv preprint arXiv:2311.07919. Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Mar- cel Blistein, Ori Ram, Dan Zhang, Evan Rosen, and 1 others. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261. Soham Deshmukh, Benjamin Elizalde, Rita Singh, and Huaming Wang. 2023. Pengi: An audio language model for audio tasks. Advances in Neural Informa- tion Processing Systems, 36:18090–18108. Ding Ding, Zeqian Ju, Yichong Leng, Songxiang Liu, Tong Liu, Zeyu Shang, Kai Shen, Wei Song, Xu Tan, Heyi Tang, and 1 others. 2025. Kimi-audio technical report. arXiv preprint arXiv:2504.18425. Qingkai Fang, Yan Zhou, Shoutao Guo, Shaolei Zhang, and Yang Feng. 2025. Llama-omni2: Llm-based real- time spoken chatbot with autoregressive streaming speech synthesis. arXiv preprint arXiv:2505.02625. Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra. 2021. Fsd50k: an open dataset of human-labeled sound events. IEEE/ACM Transactions on Audio, Speech, and Language Pro- cessing, 30:829–852. Rita Frieske and Bertram E Shi. 2024. Hallucinations in neural automatic speech recognition: Identify- ing errors and hallucinatory models. arXiv preprint arXiv:2401.01572. Sreyan Ghosh, Arushi Goel, Lasha Koroshinadze, Sang- gil Lee, Zhifeng Kong, Joao Felipe Santos, Ramani Duraiswami, Dinesh Manocha, Wei Ping, Moham- mad Shoeybi, and 1 others. 2025. Music flamingo: Scaling music understanding in audio language mod- els. arXiv preprint arXiv:2511.10289. Arushi Goel, Sreyan Ghosh, Jaehyeon Kim, Sonal Ku- mar, Zhifeng Kong, Sang-gil Lee, Chao-Han Huck Yang, Ramani Duraiswami, Dinesh Manocha, Rafael Valle, and 1 others. 2025. Audio flamingo 3: Advanc- ing audio intelligence with fully open large audio language models. arXiv preprint arXiv:2507.08128. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al- Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, and 1 others. 2024. Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14375–14385. Qingqing Huang, Aren Jansen, Joonseok Lee, Ravi Ganti, Judith Yue Li, and Daniel PW Ellis. 2022. Mu- lan: A joint embedding of music audio and natural language. In 23rd International Society for Music In- formation Retrieval Conference, ISMIR 2022, pages 559–566. International Society for Music Informa- tion Retrieval. Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, and 1 others. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276. Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle, and Bryan Catanzaro. 2024. Audio flamingo: A novel audio language model with few- shot learning and dialogue abilities. In International Conference on Machine Learning, pages 25125– 25148. PMLR. Chun-Yi Kuan, Wei-Ping Huang, and Hung-yi Lee. 2024. Understanding sounds, missing the questions: The challenge of object hallucination in large audio- language models. arXiv preprint arXiv:2406.08402. Chun-Yi Kuan and Hung-yi Lee. 2025. Can large audio- language models truly hear? tackling hallucinations with multi-task assessment and stepwise audio rea- soning. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Process- ing (ICASSP), pages 1–5. IEEE. Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2023. Halueval: A large- scale hallucination evaluation benchmark for large language models. In Proceedings of the 2023 Con- ference on Empirical Methods in Natural Language Processing, pages 6449–6464. Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Truthfulqa: Measuring how models mimic human falsehoods. In Proceedings of the 60th annual meet- ing of the association for computational linguistics (volume 1: long papers), pages 3214–3252. Bob L Sturm. 2012. An analysis of the gtzan music genre dataset. In Proceedings of the second interna- tional ACM workshop on Music information retrieval with user-centered and multimodal strategies, pages 7–12. Kim Sung-Bin, Oh Hyun-Bin, JungMok Lee, Arda Senocak, Joon Son Chung, and Tae-Hyun Oh. 2025. AVHBench: A cross-modal hallucination benchmark for audio-visual large language models. In The Thir- teenth International Conference on Learning Repre- sentations. Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun MA, and Chao Zhang. 2024. SALMONN: Towards generic hearing abilities for large language models. In The Twelfth International Conference on Learning Representa- tions. Joseph Turian, Jordie Shier, Humair Raj Khan, Bhiksha Raj, Björn W Schuller, Christian J Steinmetz, Colin Malloy, George Tzanetakis, Gissel Velarde, Kirk Mc- Nally, and 1 others. 2022. Hear: Holistic evaluation of audio representations. In NeurIPS 2021 Compe- titions and Demonstrations Track, pages 125–145. PMLR. Boyong Wu, Chao Yan, Chen Hu, Cheng Yi, Chengli Feng, Fei Tian, Feiyu Shen, Gang Yu, Haoyang Zhang, Jingbei Li, and 1 others. 2025. Step-audio 2 technical report. arXiv preprint arXiv:2507.16632. Xiyang Wu, Tianrui Guan, Dianqi Li, Shuaiyi Huang, Xiaoyu Liu, Xijun Wang, Ruiqi Xian, Abhinav Shri- vastava, Furong Huang, Jordan Boyd-Graber, and 1 others. 2024. Autohallusion: Automatic genera- tion of hallucination benchmarks for vision-language models. In Findings of the Association for Computa- tional Linguistics: EMNLP 2024, pages 8395–8419. LLM-Core-Team Xiaomi. 2025. Mimo-audio: Audio language models are few-shot learners. Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, and 1 others. 2025. Qwen2.5-omni techni- cal report. arXiv preprint arXiv:2503.20215. Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu, Ziyue Jiang, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao, Chang Zhou, and 1 others. 2024. Air- bench: Benchmarking large audio-language models via generative comprehension. In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 1979–1998. A Dataset Details A.1 Annotator Details HalluAudio is constructed through a rigor- ous human-in-the-loop annotation and validation pipeline designed to ensure label reliability at scale. Candidate QA pairs are first generated automati- cally via programmatic rules and filtering proce- dures, after which stratified subsets are manually inspected by annotators with experience in speech and audio understanding. Each instance is reviewed across multiple independent passes to verify au- dio–question consistency, ground-truth correctness, and the validity of hallucination-triggering condi- tions. To quantify annotation reliability, we compute inter-annotator agreement using Cohen’sκ. The overall agreement is 0.91, with domain-level agree- ment above 0.89 across speech, sound, and music subsets. Approximately 4–5% of candidate QA pairs were revised or discarded due to annotator disagreement. Final labels are determined through majority agreement with additional review by a senior annotator when necessary. A.2 Dataset Construction Pipeline FieldDescription QuestionThe natural language question gener- ated from a task-specific prompt tem- plate, describing the perceptual or se- mantic judgment to be made based on the input audio. AnswerThe ground-truth answer associated with the question. For binary tasks, the answer isYesorNo, while other tasks use predefined task-specific an- swer spaces. FnameThe unique identifier or filename of the audio clip used to construct the task instance, enabling traceability to the original audio source. Audio The raw audio content corresponding to the task instance, stored as serial- ized audio bytes and provided directly to the model during inference. Table 7: Data instance structure in the HalluAudio dataset. The HalluAudio dataset is constructed through a unified and systematic pipeline that transforms curated audio recordings into structured audio– question pairs for hallucination evaluation. For each selected audio clip, we generate multiple task instances by pairing the audio with parameterized prompt templates corresponding to different per- ceptual and reasoning tasks. Each dataset instance follows a consistent struc- ture, consisting of a natural language question, its ground-truth answer, a reference to the source au- dio file, and the raw audio content itself. This design ensures that all model inputs are grounded directly in the audio signal rather than intermediate representations or extracted features. Table 7 sum- marizes the fields contained in each data instance. Task instances are created by instantiating tem- plate variables using verified annotations associ- ated with each audio clip, such as transcriptions, sound event labels, or musical attributes. Valid queries are guaranteed to be answerable from the audio evidence alone, while invalid or unanswer- able queries are intentionally constructed by ref- erencing attributes absent from the audio. All in- stances are automatically validated to ensure con- sistency between the question, answer space, and audio content. A.3 Prompt Template Design Prompt templates in HalluAudio are designed to generate theQuestionfield of each dataset in- stance, probing hallucination behaviors under di- verse linguistic and perceptual conditions while preserving consistent task semantics. For each task category, multiple templates are constructed with varied phrasing styles and contextual formulations to reduce sensitivity to prompt-specific artifacts. In the speech domain, prompt templates target fine-grained acoustic and semantic reasoning, in- cluding word occurrence, temporal ordering, count- ing, and speaker-related attributes. In the environ- mental sound domain, templates focus on sound event perception and co-occurrence reasoning, such as sound presence, overlap detection, and multi- label identification. In the music domain, prompts emphasize perceptual and categorical judgments, including instrument identification, genre match- ing, loudness comparison, and speech–music dis- crimination. Across all domains, both valid and inten- tionally invalid prompt templates are included. Valid prompts correspond to answerable ques- tions grounded in the audio content, while invalid prompts are constructed to be unanswerable by de- sign, serving as controlled probes for hallucinated responses and refusal behaviors. All templates are parameterized and instantiated automatically, ensuring scalable and reproducible generation of Figure 5: Yes/No Bias Analysis in music domain. From left to right: Yes-pred Ratio, Unrelated Error Ratio, and Conditional Accuracy. Higher Yes-pred Ratios combined with low conditional accuracy indicate strong affirmative bias rather than evidence-grounded binary reasoning. question–audio pairs. B Music data analysis B.1 Yes/No Bias Analysis In Figure 5, the Yes/No bias analysis reveals pro- nounced asymmetry across task types and mod- els.The Yes prediction ratio shows that sev- eral models, notably Qwen2.5-Omni and GPT-4o- Audio, exhibit near-saturated affirmative responses on ordering-related queries, indicating a strong ten- dency to default to positive judgments when reason- ing about musical structure. Genre and instrument identification tasks display more dispersed behav- ior, suggesting comparatively better calibration un- der high-level semantic cues, while matching tasks expose substantial variability across models. The unrelated error ratio highlights that high affirmative bias does not necessarily translate to semantically grounded errors. Ordering tasks con- sistently induce the highest unrelated responses, implying that structural musical reasoning is par- ticularly vulnerable to hallucinated content rather than simple misclassification. In contrast, genre- related queries show lower unrelated ratios, reflect- ing stronger alignment between model responses and the intended decision space. Finally, conditional accuracy exposes a clear im- balance between positive and negative cases. Mod- els with strong affirmative tendencies achieve high accuracy on positive instances but degrade sharply on negative matching tasks, revealing limited sen- sitivity to absence or contradiction in musical evi- dence. Overall, these results indicate that hallucina- tion in music-centric tasks is tightly coupled with structural reasoning demands, and that high-level semantic understanding alone is insufficient to en- sure reliable binary judgment in complex musical contexts. B.2 False Refusal Analysis. In Figure 6, false refusals are sparse overall but highly task-specific: Qwen2-Audio collapses on order and loudness, while Gemini-2.5-Flash shows elevated refusal across genre, counting, and match- ing tasks. Figure 6: False Refusal Rate in music domain. Higher values indicate a stronger tendency to refuse answering despite the presence of a valid ground-truth answer. C Post-hoc Robustness Test ModelOrig.Para.Diff. Qwen2-Audio-7B50.149.90.2 Kimi-Audio-7B48.447.90.5 Pengi22.720.91.8 MiMo-Audio-7B-Instruct61.360.80.5 Step-Audio-2-mini56.856.50.3 Average47.947.20.7 Table 8: Robustness(%) of different audio-language models under semantic paraphrasing. Results are aver- aged over 1,000 sampled QA pairs across all domains. To evaluate potential linguistic bias introduced by template-based prompt generation, we con- duct a post-hoc robustness analysis using semanti- cally equivalent paraphrases across multiple audio- language models. We randomly sample 1,000 QA pairs from Hal- luAudio, covering speech, environmental sound, and music domains. For each instance, we manu- ally construct a paraphrased variant that preserves semantic meaning while altering surface wording and syntactic structure. We evaluate both original and paraphrased prompts on five representative models, including Qwen2-Audio-7B, Kimi-Audio-7B, Pengi, MiMo- Audio-7B-Instruct, and Step-Audio-2-mini, under identical inference settings. As shown in Table 8, all models exhibit minimal performance variation under paraphrasing. The absolute difference re- mains consistently small, ranging from 0.2% to 1.8%, with an average deviation of 0.7%. This consistency across diverse architectures in- dicates that model behavior is largely invariant to superficial linguistic variations. The results suggest that performance differences are driven primarily by task structure and acoustic reasoning require- ments rather than prompt wording. Overall, these findings provide strong evidence that the template-based generation strategy in Hal- luAudio does not introduce significant linguistic bias, and that the benchmark reliably reflects hal- lucination behavior instead of prompt sensitivity artifacts. D LALMs Participating in Benchmark Based on the different focuses of different LALMs, we selected different LALMs for different domains, and the detailed data is listed in Table 9. Qwen-Audio-Chat: An open-source multi- modal audio-language model by Alibaba Cloud. It extends the Qwen-Audio foundation by instruc- tion fine-tuning to support multi-turn spoken di- alogue. The model accepts diverse audio inputs along with text and generates text responses. Qwen- Audio-Chat is designed for comprehensive audio understanding – including speech reasoning, sound classification, music appreciation, and even audio- based editing – within conversational contexts. Qwen2-Audio-7B: A 7-billion-parameter audio- aware LLM from Alibaba’s Qwen series. It sup- ports two interactive modes: "voice chat" and "au- dio analysis". Qwen2-Audio-7B can perform tasks like ASR, audio classification, speech-to-text trans- ModelSpeech Sound Music Source Qwen-Audio-Chat✓Open Qwen2-Audio-7B✓Open Qwen2.5-Omni-7B✓Open Llama-3.1-8B-Omni✓Open Llama-Omni2-7B✓Open Kimi-Audio-7B✓Open Phi-4-Multimodal✓Open Audio Flamingo-3✓Open Music Flamingo✓Open Pengi✓Open MiMo-Audio-7B-Instruct✓Open Step-Audio-2-mini✓Open GPT-4o-Audio-Preview✓Closed Gemini-2.5-Flash✓Closed Table 9: LALMs evaluated in HalluAudio across differ- ent domains. lation, and emotion or sound recognition across multiple languages. The system is released with an instruct-tuned variant and achieves strong bench- marks on standard speech and audio understanding tasks. Qwen2.5-Omni-7B: A 7B open-source multi- modal model by Alibaba. It is designed to perceive and integrate text, images, audio, and video in- puts simultaneously, while generating textual and natural speech outputs in real time. Qwen2.5- Omni uses a "Thinker-Talker" architecture with a time-aligned embedding scheme to synchronize audio/video timestamps. This enables features like real-time voice and video chat. In evaluations, Qwen2.5-Omni outperforms similarly-sized single- modality models on joint tasks and exceeds the audio capabilities of Qwen2-Audio. LLaMA-3.1-8B-Omni: A speech-enabled LLM built on Meta’s Llama-3.1 8B Instruct model. LLaMA-Omni integrates a pretrained speech en- coder, a speech adaptor, and a streaming speech decoder with the base LLM. This design eliminates the need for intermediate transcription: it directly generates text and speech responses from spoken instructions. The result is low-latency, high-quality spoken dialogue – the model can answer and even speak back with a latency on the order of a few hundred milliseconds while maintaining content fidelity and natural style. LLaMA-Omni2-7B: A 7B variant of the LLaMA-Omni2 series. Built on Qwen2.5, this model incorporates a speech encoder and an autore- gressive streaming speech decoder into the LLM. It enables real-time spoken chat: given speech in- put, the model can generate text or speech answers on the fly. Even though it was trained on only ̃ 200K multi-turn speech QA pairs, LLaMA-Omni2 models exhibit strong performance on spoken dia- log and instruction tasks, surpassing prior speech- language models on several benchmarks. Kimi-Audio-7B: A 7B open-source audio foun- dation model by MoonshotAI. Kimi-Audio is de- signed to handle a wide variety of audio tasks in one model. Its input encoder is "hybrid": incom- ing audio is tokenized into discrete semantic to- kens and also encoded into continuous acoustic features. These features feed into a transformer LLM core, which has parallel output heads for gen- erating text tokens and audio tokens. Kimi-Audio was pretrained on over 13 million hours of diverse audio and text, giving it strong audio-language un- derstanding. It achieves state-of-the-art results on tasks like ASR, audio question-answering, audio captioning, emotion recognition, sound classifica- tion, and even end-to-end speech conversation. Phi-4-Multimodal: A 3.8B open-source small multimodal model from Microsoft.It unifies text, vision, and speech/audio in one model us- ing a "mixture-of-LoRAs" approach: the base lan- guage model is frozen and modality-specific LoRA adapters are added for vision and audio. Phi- 4-Multimodal thus natively supports inputs like speech+text or image+audio. Despite its compact size, it achieves very strong performance on speech tasks for a model of its scale: for example, it ranks first on the open multilingual ASR leaderboard and excels at speech translation and QA. Notably, it is the first open-source model to include a speech summarization capability, highlighting its compre- hensive audio understanding. Audio Flamingo 3: An open-source Large Audio-Language Model by NVIDIA. AF3 ad- vances unified audio reasoning across speech, en- vironmental sounds, and music.It builds on Flamingo-style architecture with a unified audio en- coder and adds novel features like flexible chain-of- thought reasoning and long-context comprehension. AF3 supports multi-turn audio dialogues and even voice-to-voice conversational response. In bench- marks, Audio Flamingo 3 sets new state-of-the-art scores on over 20 public audio understanding and reasoning tasks. Music Flamingo: An open-source NVIDIA model specialized for music understanding. Music Flamingo analyzes complex musical audio with deep musical knowledge. It can generate rich, theory-aware captions and answers about music attributes. The model is trained with reasoning- centric methods and can process full-length songs. In evaluations it establishes new SOTA on more than 10 music-related tasks. Pengi: An audio-language model by Microsoft. Pengi reframes all audio tasks as text-generation tasks. It uses an audio encoder to convert any input audio into embeddings, concatenates this with any text prompt, and feeds it into a pretrained frozen language model. This unified approach allows both open-ended tasks and closed tasks to be handled without task-specific fine-tuning. MiMo-Audio-7B-Instruct: A 7B open-source audio LLM by Xiaomi. The core MiMo-Audio model is pretrained on a massive scale, which enables emergent few-shot generalization to new audio tasks. Even without fine-tuning, the 7B base MiMo-Audio already achieves state-of-the- art open-model performance on standard speech and audio understanding benchmarks. After in- struction fine-tuning, MiMo-Audio yields leading open-source results on audio comprehension, spo- ken dialogue, and TTS instructions. It can gener- alize to tasks not in its training data and produce realistic continuous speech as part of its outputs. Step-Audio-2-Mini: The 8B-parameter variant of StepFun’s Step-Audio 2 family. Step-Audio 2 is an end-to-end multimodal audio-language system designed for "industry-strength" audio understand- ing and conversation. It integrates a latent audio encoder and uses reinforcement learning focused on reasoning. Importantly, Step-Audio-2 generates discrete audio tokens as part of its output, allowing it to capture paralinguistic cues in its responses. It also incorporates retrieval to ground its knowledge and reduce hallucinations. Trained on millions of hours of speech/audio, Step-Audio 2 achieves state-of-the-art performance on diverse audio un- derstanding and conversational benchmarks. GPT-4o-Audio-Preview:A closed-source audio-capable model from OpenAI. In preview re- lease, GPT-4o-Audio accepts both text and audio as input and can produce either text or audio out- puts. According to OpenAI’s API documentation, it supports a very large context window (128,000 tokens) for audio/text and is accessed via the Chat Completions endpoint. Gemini-2.5-Flash: Google’s proprietary audio LLM optimized for real-time voice interactions. The latest "Flash" version improves instruction- following and conversation smoothness for live voice agents. Gemini-2.5-Flash Native Audio can interpret complex spoken instructions, trigger ex- ternal tools or function calls, and maintain natu- ral multi-turn dialogue. Google has deployed it in products like Google Translate and in Google AI/Vertex AI for building voice agents. It supports 70+ languages in live translation, enabling seam- less voice-based communication across languages. E Output Normalization Due to the open-ended generation nature of LALMs, raw model outputs exhibit substantial vari- ability in surface forms. For example, binary deci- sions may be expressed as concise tokens, extended explanations, or embedded within free-form reason- ing. Such diversity poses challenges for consistent and fair evaluation across models and tasks. To ensure reliable metric computation, we apply a text-level output normalization procedure prior to labeling. The normalization process standardizes model responses while preserving their semantic intent. Specifically, we first convert all outputs to lowercase and remove punctuation and redundant whitespace. For binary decision tasks, normalized outputs are mapped to canonical labels (YesorNo) via keyword matching. For counting tasks, numer- ical values are extracted from textual responses when present. In addition, responses indicating in- ability to answer or lack of access to audio content are explicitly mapped to a unified Refusal label. Table 10 illustrates representative examples of raw model outputs and their corresponding normal- ized forms. Output TypeRaw Model OutputNormalized Form BinaryYes, it does.Yes BinaryNo, it is absent.No CountIt appears three times.3 RefusalI cannot determine.Refusal Table 10: Examples of output normalization across dif- ferent response types. F Representative Failure Cases For each category, we report qualitative examples where multiple LALMs are evaluated on the same audio–question pair, highlighting systematic hal- lucination behaviors that are not fully captured by aggregate metrics. F.1 Hallucinated Affirmative Responses This category captures cases where models pro- duce confident affirmative or content-bearing re- sponses despite the absence of supporting acoustic evidence. As illustrated in Figure 7, when pre- sented with an audio segment containing no in- telligible speech, many LALMs nonetheless gen- erate fluent and specific transcriptions, ranging from plausible everyday utterances to long-form, domain-irrelevant statements. These responses are not random noise but linguistically well-formed, indicating that models default to language priors rather than grounding their outputs in the actual audio input. Such hallucinated affirmative behaviors are es- pecially prevalent in open-ended recognition and description tasks, where the absence of explicit negative supervision encourages over-generation. While some models appropriately abstain or ac- knowledge uncertainty, others fabricate detailed content with high confidence, revealing a system- atic affirmative bias under perceptual uncertainty. This failure mode highlights a critical reliability gap in current LALMs: when acoustic evidence is weak or missing, models often prefer generat- ing a plausible answer over producing a grounded refusal or null response. F.2 Incorrect Task-grounded Answers In this category, models attempt to perform the instructed task, such as detecting overlapping speech, identifying speech content, or making binary/multiple-choice decisions, but the answers they produce are inconsistent with the actual audio content. Unlike hallucinated affirmatives, where models fabricate content in the absence of evidence, responses in this category remain task-relevant yet are factually incorrect. These errors often occur in tasks with binary or limited-choice outputs, where subtle audio cues or temporal patterns must be interpreted correctly. Common causes include overgeneralization, mis- alignment between model predictions and the au- dio timeline, or reliance on superficial heuristics instead of the underlying auditory evidence. As illustrated in Figure 8, some models confidently produce answers that contradict the ground truth, while others hedge, partially refusing to commit, leading to a mix of incorrect, incomplete, or am- biguous task-grounded outputs. Overall, this phenomenon highlights that even when models understand the task instruction, achieving factual accuracy in audio-based reason- ing remains a significant challenge. Case 1: Speech – Recognition Prompt: What does the speaker say in the recording? Reference Answer: None. Erroneous Model Outputs: •Qwen-Audio-Chat: The speaker says in English: “You can see the logic in the data and the model is quite good.” • Qwen2-Audio-7B: The speaker says “Kids are talking by the door.” • LLaMA-3.1-8B-Omni: Yes, I can help you with that. What would you like to know about the weather? • LLaMA-Omni2-7B: I’m sorry, I don’t have access to the recording. • Kimi-Audio-7B: I don’t know. • Phi-4-Multimodal: The speaker says “and subscribe for more videos!”. • Pengi: static. • MiMo-Audio-7B-Instruct: The speaker says “Thank you for watching!”. • Step-Audio-2-Mini: There is white noise. (Correct) • GPT-4o-Audio-Preview: “Currently, higher education is facing a crisis with rising tuition costs. . . ” • Gemini-2.5-Flash: “She asked if I had seen the movie.” This case demonstrates a semantic hallucination, where models fabricate linguistic content despite the absence of intelligible speech in the audio. Figure 7: Semantic hallucination examples in speech recognition tasks. F.3 False Refusal Responses This category covers cases where models refuse to answer questions that are, in fact, answerable given the provided audio input. Instead of ana- lyzing the audio content and producing a valid response, models default to refusal behaviors, of- ten expressed through generic uncertainty state- ments, access-related disclaimers, or requests for additional context. Such false refusals typically reflect overly con- servative assumptions about model capabilities or safety constraints, rather than genuine ambiguity in the audio signal. Notably, these errors frequently arise in tasks that require recognizing the absence of a target event or attribute, where the correct an- swer is negative but still well-defined. As a result, models avoid committing to an answer despite suf- ficient evidence being available in the audio. F.4 Cross-domain Failure Patterns Beyond individual examples, we observe clear domain-dependent trends in model failure behav- iors. For speech-related tasks, models frequently produce hallucinated affirmatives or incorrect task- grounded answers, particularly in questions involv- ing word order, counting, or temporal relations. These errors suggest that models often rely on coarse linguistic priors or overgeneralized patterns rather than precise alignment with the underlying audio timeline. In contrast, environmental sound tasks tend to elicit false refusals or overly cautious responses. Even when the acoustic evidence is clear, mod- els may default to generic uncertainty or access- related disclaimers, indicating conservative as- sumptions about input availability or ambiguity in non-linguistic audio signals. Case 2: Speech – Temporal Comparison Prompt: Do the two speakers’ voices overlap in the recording? Reference Answer: Yes. Erroneous Model Outputs: • Qwen-Audio-Chat: No, the two speakers’ voices do not overlap in the recording. •Qwen2-Audio-7B: Yes, the two speakers’ voices overlap in the recording as they both speak from time 0.00 to 2.39 and then from time 2.47 to 4.95. (Correct) • LLaMA-3.1-8B-Omni: No, she did not live with her grandfather. • LLaMA-Omni2-7B: Yes, the two speakers’ voices do overlap in the recording. (Correct) • Kimi-Audio-7B: no. • Phi-4-Multimodal: The audio clip does not provide enough information to determine if the two speakers’ voices overlap. • Pengi: yes. (Correct) • MiMo-Audio-7B-Instruct: Yes. (Correct) • Step-Audio-2-Mini: No, the two speakers’ voices do not overlap in the recording. • GPT-4o-Audio-Preview: No, the two speakers’ voices do not overlap in the recording. •Gemini-2.5-Flash: As an AI, I don’t have access to external audio files to listen to them. Therefore, I cannot tell you if the two speakers’ voices overlap in the recording. This case exemplifies an incorrect task-grounded answer, where models attempt to perform the instructed task but produce answers that contradict the actual audio content. Unlike hallucinated affirmatives, the responses are relevant to the task but factually inaccurate. Such errors are particu- larly common in binary or multiple-choice classification tasks, where the model may overgeneralize, misinterpret subtle temporal cues, or rely on heuristics that do not align with the audio evidence. Figure 8: Incorrect task-grounded responses in a temporal comparison task, illustrating models’ tendencies to produce task-relevant but factually inconsistent outputs. Music-related tasks exhibit a distinct failure pro- file, characterized by label confusion, task devi- ation, and conservative abstention. Models fre- quently struggle to distinguish structurally similar instruments or sound events, leading them to shift attention to irrelevant acoustic attributes or to avoid making definitive judgments. This behavior high- lights persistent challenges in fine-grained acoustic reasoning and semantic grounding within complex musical contexts. Case 3: Speech – Recognition Prompt: What does the female say in the recording? Reference Answer: No female is speaking in the recording. Erroneous Model Outputs: • Qwen2-Audio-7B: I’m sorry, I can’t answer this question. •LLaMA-3.1-8B-Omni: I’m happy to help! However, I didn’t quite catch what you’re referring to. Could you please provide more context or clarify your question? •LLaMA-Omni2-7B: I’m not aware of any specific recording that matches your description. Could you provide more context or information about the recording you’re referring to? • Phi-4-Multimodal: I’m sorry, I cannot provide an answer as the audio clip is not provided. •GPT-4o-Audio-Preview: I’m sorry, but I can’t help with identifying specific speakers from audio recordings. If you need help with something else, feel free to ask! •Gemini-2.5-Flash: I’m sorry, but I need the audio recording to tell you what the female says. Please provide the audio. This case exemplifies a false refusal response, where the question is answerable given the audio input, yet the models incorrectly refuse to provide an answer. Although the correct response requires recognizing the absence of a female speaker, several models default to generic uncertainty or access- related disclaimers, reflecting overly conservative refusal behavior rather than genuine ambiguity in the audio. Cross-domain Contrast: Word Order Model: Qwen2-Audio-7B • Speech Prompt: Does the word "and" appear before the word "used" in the recording? Reference: No. Model Output: Yes, the word "and" appears before the word "used" in the recording. Issue: Incorrect task-grounded affirmative, where the model confidently answers the word-order question but contradicts the actual temporal structure of the audio. • Environmental Sound Prompt: Does "Gasp" appear before "Clapping" in the recording? Reference: Yes. Model Output: No, the "Gasp" sound appears after the "Clapping" sound in the recording. Issue: Incorrect task-grounded answer in temporal ordering, indicating a failure to correctly infer event sequence despite clear acoustic evidence. • Music Prompt: Does "ding" appear before "ding-dong" in the recording? Reference: Yes. Model Output: I’m not sure which sound is louder just from the audio. You could try listening more closely or maybe compare the volume levels on your device. If you can tell me more about the sounds, like what they are or how they were played, that might help. So, what do you think? Issue: Task deviation caused by label confusion, where the model shifts from temporal compari- son to an unrelated acoustic attribute, resulting in an off-task and non-informative response. This example demonstrates that the same LALM exhibits distinct failure modes across domains when performing structurally similar word-order or temporal comparison tasks, ranging from incorrect affirmatives to task deviation and ambiguous responses.