Paper deep dive
VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories
Qian Zhang, Yuqin Cao, Yixuan Gao, Xiongkuo Min
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/14/2026, 2:21:52 AM
Summary
VidAudio-Bench is a comprehensive multi-task benchmark designed to evaluate Video-to-Audio (V2A) and Video-Text-to-Audio (VT2A) generation models. It covers four audio categories (sound effects, music, speech, and singing) using 1,634 video-text pairs and 13 fine-grained, reference-free metrics. The study reveals that current models struggle with speech and singing compared to sound effects and highlights a trade-off in VT2A models between visual grounding and instruction following.
Entities (5)
Relation Signals (3)
VidAudio-Bench â evaluates â V2A
confidence 100% ¡ we propose VidAudio-Bench, a multi-task benchmark for V2A evaluation
VidAudio-Bench â evaluates â VT2A
confidence 100% ¡ VidAudio-Bench, the first comprehensive multi-task benchmark designed for both V2A and VT2A evaluation
Qwen3-VL â supports â VT2A
confidence 90% ¡ we utilize Qwen3-VL to extract descriptions... enabling a cleaner test of visual grounding
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Video-to-Audio (V2A) generation is essential for immersive multimedia experiences, yet its evaluation remains underexplored. Existing benchmarks typically assess diverse audio types under a unified protocol, overlooking the fine-grained requirements of distinct audio categories. To address this gap, we propose VidAudio-Bench, a multi-task benchmark for V2A evaluation with four key features: (1) Broad Coverage: It encompasses four representative audio categories - sound effects, music, speech, and singing - under both V2A and Video-Text-to-Audio (VT2A) settings. (2) Extensive Evaluation: It comprises 1,634 video-text pairs and benchmarks 11 state-of-the-art generation models. (3) Comprehensive Metrics: It introduces 13 task-specific, reference-free metrics to systematically assess audio quality, video-audio consistency, and text-audio consistency. (4) Human Alignment: It validates all metrics through subjective studies, demonstrating strong consistency with human preferences. Experimental results reveal that current V2A models perform poorly in speech and singing compared to sound effects. Our VT2A results further highlight a fundamental tension between instruction following and visually grounded generation: stronger visual conditioning improves video-audio alignment, but often at the cost of generating the intended audio category. These findings establish VidAudio-Bench as a comprehensive and scalable framework for diagnosing V2A systems and provide new insights into multimodal audio generation.
Tags
Links
- Source: https://arxiv.org/abs/2604.10542v1
- Canonical: https://arxiv.org/abs/2604.10542v1
Trouble viewing inline? Open PDF directly â
Full Text
93,260 characters extracted from source content.
Expand or collapse full text
VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories Qian Zhang Shanghai Jiaotong University Shanghai, China zq1729@sjtu.edu.cn Yuqin Cao Shanghai Jiaotong University Shanghai, China caoyuqin@sjtu.edu.cn Yixuan Gao Shanghai Jiaotong University Shanghai, China gaoyixuan@sjtu.edu.cn Xiongkuo Min Shanghai Jiaotong University Shanghai, China minxiongkuo@sjtu.edu.cn S o u n d E f f e c t s M u s i c S p e e c h S i n g i n g Generated Audios Video-to-Audio Video-Text-to-Audio ... Evaluation Dimension Suite Generated Captions Realistic foley sound of a lion roaring with its mouth open on grass. Video-Audio Consistency Audio Quality Text-Audio Consistency Eight for music Six for SFX Nine for speech Ten for singing Human Preference Study Alignment Verification VidAudio-Bench 1.6k+ videos Fidelity Perception Musicality Affective Alignment Semantic Alignment Instruction Following Rhythmic Sync Temperal Sync Identity Consistency Lip Sync Aesthetic Semantic Correspondence Intelligibility Figure 1: Overview of VidAudio-Bench. We categorize audio generation into four task types: sound effects, music, speech, and singing. We further introduce two input paradigms, V2A and VT2A, to analyze how adding textual descriptions changes generation behavior. Our evaluation suite spans Audio Quality, Video-Audio Consistency, and Text-Audio Consistency, covering 13 fine-grained dimensions. Human preference studies show strong correlation between our metrics and human perception. Abstract Video-to-Audio (V2A) generation is essential for immersive multi- media experiences, yet its evaluation remains underexplored. Ex- isting benchmarks typically assess diverse audio types under a unified protocol, overlooking the fine-grained requirements of dis- tinct audio categories. To address this gap, we propose VidAudio- Bench, a multi-task benchmark for V2A evaluation with four key features: (1) Broad Coverage: It encompasses four representative audio categoriesâsound effects, music, speech, and singingâunder both V2A and Video-Text-to-Audio (VT2A) settings. (2) Extensive Evaluation: It comprises 1,634 video-text pairs and benchmarks 11 state-of-the-art generation models. (3) Comprehensive Metrics: It introduces 13 task-specific, reference-free metrics to systematically assess audio quality, videoâaudio consistency, and textâaudio con- sistency. (4) Human Alignment: It validates all metrics through subjective studies, demonstrating strong consistency with human preferences. Experimental results reveal that current V2A models perform poorly in speech and singing compared to sound effects. Our VT2A results further highlight a fundamental tension between instruction following and visually grounded generation: stronger visual conditioning improves video-audio alignment, but often at the cost of generating the intended audio category. These findings establish VidAudio-Bench as a comprehensive and scalable frame- work for diagnosing V2A systems and provide new insights into multimodal audio generation. Keywords Video-to-Audio Benchmark, Multimodal Evaluation, Cross-Modal Alignment, Audio Quality Assessment, Perceptual Quality 1 Introduction While audio generation has made significant progress through Text-to-Audio (T2A) [20,27,34,62,63,71] and Image-to-Audio (I2A) [7,23,40] models, relying solely on static or textual inputs 1 arXiv:2604.10542v1 [cs.SD] 12 Apr 2026 often fails to capture precise temporal dynamics, complex spa- tial environments, and realistic physical interactions. This limi- tation has driven the shift toward Video-to-Audio (V2A) genera- tion [26,33,50,58,59,68], where rich visual cues provide explicit conditions for dynamic audio synthesis. Recent V2A models have evolved toward multimodal and multi-task frameworks [12,51,61]. Pioneering models such as AudioX [45] and AudioGen-Omni [52] enable the flexible generation of diverse audio types, emphasiz- ing the importance of not just sound effects, but also music and speech. Similarly, large-scale frameworks like Kling-Foley [51] and Audiobox-Aesthetics [47] categorize their training data into distinct modalities (e.g., sound effects, music, speech and singing). These advances call for more fine-grained evaluation protocols that can reflect the specific requirements of different audio categories. However, existing evaluation methodologies remain monolithic and outdated. First, current benchmarks suffer from the "one-size- fits-all" pitfall. Despite the distinct acoustic and semantic properties of different audio tasks, generative models are almost exclusively evaluated using generic distribution-level metrics, such as FrĂŠchet Audio Distance (FAD) [37], Kullback-Leibler Divergence (KL), Incep- tion Score (IS) [38], and cross-modal similarity (e.g., ImageBind [17]). These metrics fail to capture task-dependent requirements, such as the intelligibility and lip-sync precision necessary for speech, or the melodic coherence required for music. Second, the eval- uation data itself is problematic. Most models still evaluate on the raw VGG-Sound test set [6], which, as explicitly noted by the VGGSounder [73], suffers from severe limitations including co- occurring classes, overlapping sounds, and modality misalignment. A further critical bottleneck in current V2A research lies in the ambiguous role of text. Early models such as FRIEREN [53] rely solely on visual inputs, whereas most recent architectures naturally support joint videoâtext conditioning (VT2A) [31,39,45,51,61]. However, during evaluation, it remains unclear whether the gener- ated audio is successfully grounded in the visual content, or if the model mainly follows explicit acoustic textual prompts (e.g., "the sound of a dog barking") as a shortcut. This calls for an evaluation paradigm that can disentangle modelsâ visual understanding capa- bility from auxiliary text guidance, which is particularly vital for context-aware multimedia applications. To address these limitations, we propose VidAudio-Bench, the first comprehensive multi-task benchmark designed for both V2A and VT2A evaluation. As shown in Figure 1, VidAudio-Bench en- compasses four representative audio categories: sound effects (SFX), music, speech, and singing. For each category, we construct care- fully curated evaluation subsets, totaling 1,634 video-text pairs with strong audio-visual correlation. To move beyond a monolithic evaluation, we introduce a suite of task-specific, reference-free evaluation protocols. VidAudio-Bench comprises 13 fine-grained dimensions covering audio quality, cross-modal alignment, and domain-specific attributes. Leveraging recent advances in multi- modal large language models (MLLMs) [10,22,41,60,66], which have been shown to be effective for evaluation [29], we incorpo- rate an MLLM-as-a-Judge framework for three of our advanced dimensions. Crucially, comprehensive user studies confirm that our evaluation suite is closely aligned with human perception. Moreover, VidAudio-Bench introduces a novel VT2A evaluation paradigm that explicitly probes a modelâs visual understanding. In- stead of providing textual prompts that dictate the target sound, we supply dense visual descriptions of the scene. This zero-information- leak design prevents models from exploiting acoustic textual short- cuts, enabling a cleaner test of visual grounding and instruction following. Our results further reveal a counterintuitive effect: dense captions often improve semantic alignment but weaken instruction following, leading to target-miss errors. In summary, our main contributions are threefold: ⢠Comprehensive Multi-Task Benchmark: We present VidAudio- Bench, which organizes video-to-audio generation into four sub- tasks: SFX, music, speech, and singing. Featuring over 400 highly correlated video-text pairs per task and 13 fine-grained dimen- sions, it provides a more systematic evaluation framework. ⢠Novel VT2A Evaluation Setting: We introduce a VT2A evalu- ation setting using dense visual descriptions instead of explicit audio prompts, enabling a cleaner assessment of visual under- standing and reducing shortcut reliance on textual acoustic cues. â˘Extensive Benchmarking and Insights: We benchmark a broad range of state-of-the-art models and show that current systems still struggle with domain-specific generation and visu- ally grounded audio synthesis. Our study further uncovers the dual role of visual prompts in VT2A generation. 2 Related Works 2.1 Audio Generation Models Video-to-Audio (V2A) aims to generate semantically aligned and temporally synchronized audio for silent videos. Early methods [69] directly mapped frames to waveforms, while later methods like SpecVQGAN [23] and Im2Wav [40] generate latent audio repre- sentations conditioned on visual features extracted by models like CLIP [36]. To overcome data scarcity, recent work leverages pre- trained T2A models [50,58,68]. For example, Seeing and Hear- ing [59] converts videos into AudioLDM [30] prompts via Im- ageBind [17], while V2A-Mapper [50] aligns visual features with CLAP [57]. To further improve temporal alignment, Diff-Foley [33] introduces contrastive audio-visual pretraining, whereas Foley- Crafter [68], ReWaS [26], and SonicVLM [58] integrate time-aware control modules. Recent efforts also focus on alignment and ef- ficiency. V-AURA [48] uses high-frame-rate visual features, and FRIEREN [53] accelerates sampling via efficient rectified flow match- ing (RFM). Furthermore, unified multimodal training (e.g., VATT [1], MMAudio [9]) has been explored to improve semantic consistency. Beyond this, some works investigate more flexible control. Multi- Foley [8] supports multimodal conditioning, and ThinkSound [31] introduces Chain-of-Thought (CoT) reasoning for guided synthesis. More recent models, such as HunyuanVideo-Foley [39] and Kling- Foley [51], adopt advanced diffusion transformer architectures to further improve audio quality and synchronization. 2.2 Evaluation Benchmarks The rapid progress of Artificial Intelligence Generated Content (AIGC) has driven the development of systematic evaluation bench- marks. Existing frameworks such as VBench [21] and EvalCrafter [32] 2 provide multi-dimensional evaluation for text-to-video (T2V) gener- ation, while TTA-Bench [49] and T2A-EpicBench [54] focus on T2A generation quality. More recently, benchmarks for joint audio-video generation have also emerged. For example, VABench [18] presents a 15-dimension framework for evaluating Text-to-Audio-Video (T2AV) and Image-to-Audio-Video (I2AV), and T2AV-Compass [5] combines objective signal-level metrics with subjective MLLM-as- a-Judge evaluation. However, existing benchmarks do not fully address the challenges of V2A evaluation. Unlike T2AV, V2A re- quires models to infer plausible audio content from visual cues alone, leading to greater ambiguity and a stronger need for task- aware assessment. In addition, although T2AV-Compass groups audio into sounds, speech, and music, its evaluation remains rela- tively unified, without incorporating the distinct criteria required by different audio categories. Consequently, V2A evaluation is still underdeveloped and fragmented, and the field lacks a standardized multi-task benchmark. To this end, we propose VidAudio-Bench, a task-specific and reference-free benchmark for evaluating V2A and VT2A systems across diverse audio categories. 3 Benchmark Construction 3.1 Task Design Motivated by recent advances in audio generation models [51,52], we divide audio generation into four representative sub-tasks based on acoustic properties, semantic functions, and cross-modal align- ment requirements. This taxonomy provides a structured frame- work for assessing the challenges of each audio category. Sound Effects (SFX). Sound effects (e.g., ambient, Foley, and in- teraction sounds) are tightly coupled with visual events. The key challenge lies in accurately recognizing the occurring events and generating sounds that are accurately synchronized with them. Music. Unlike transient sound effects, music is sustained and struc- tured over time. Depending on the visual context, this task involves two distinct challenges, leading us to define two sub-categories: â˘Instrumental Performance: Videos depicting musicians playing instruments. The generation must exhibit a frame-level align- ment between visual actions (e.g., pressing piano keys, bowing a violin) and the resulting musical notes and rhythms. â˘Background Music (BGM): Videos requiring music to support the narrative atmosphere and emotional progression. Here, the focus shifts from strict temporal synchronization to broader semantic and affective alignment with the sceneâs mood and pacing. Speech. The speech generation task focuses on synthesizing nat- ural and high-fidelity human voices from talking faces. The key challenge is to produce intelligible speech that is precisely synchro- nized with lip movements and consistent with the speakerâs visible characteristics, such as timbre, age, gender, and expression. Singing. Singing combines characteristics of both speech and music. The main challenge is to generate melodious vocals that align with both the musical rhythm and the singerâs lip movements, while preserving lyric intelligibility and vocal identity consistency. 3.2 Dataset Construction Corresponding to the four task definitions in Section 3.1, we con- struct customized evaluation subsets for each task. This section details the selection criteria, source datasets, and dataset statistics. 3.2.1 Data Selection Criteria. A fundamental prerequisite for eval- uating V2A generation is ensuring a high audio-visual correlation. Specifically, the visual content should provide sufficient informa- tion to reliably predict the associated sounds. To ensure this, we establish strict filtering criteria for each task: â˘Sound Effects: Videos must contain a visible sound source and salient motion. The sound-producing objects must be strictly on-screen, explicitly excluding any off-screen voiceovers or ambient noises without visual grounding. â˘Music: For Instrumental Performance, both the musician and instrument must be clearly visible without severe occlusion or ambiguous visual cues. For Background Music, the video must exhibit a distinct emotional mood or rhythmic cuts that naturally align with the musical narrative. â˘Speech and Singing: Videos must present a single frontal face without occlusion. The video should clearly reveal lip movements; for singing samples, it should further provide evident expressive cues, such as facial expressions or upper- body motions, to facilitate reliable evaluation. 3.2.2 Source Datasets and Subset Construction. To satisfy the re- quirements of the defined tasks, we curate suitable clips from sev- eral large-scale, high-quality public datasets. Further details on the subsets are provided in Appendix A.1. â˘VGGSounder [73]: A re-annotated multi-label dataset con- taining 15,446 clips across 309 classes with over 40,000 labels. Its strong audio-visual grounding makes it suitable for event- centric audio generation. Based on our task taxonomy, we categorize its classes and sample 400 videos for the SFX task and 191 videos for the Instrumental Performance subset. â˘HarmonySet [70]: A large-scale video-music dataset contain- ing 48,328 video-music pairs, in which background music is intentionally matched to the visual narrative. From this dataset, we extract 231 high-quality clips for the BGM task. â˘AVSpeech [15]: A large-scale dataset containing 4,700 hours of clean, single-speaker lectures and TED talks, ensuring clear correspondence between speech and visible faces. From its test set, we sample 412 clips for the Speech task. â˘Acappella [35]: Designed for multimodal singing voice sep- aration, this dataset comprises 46 hours of high-quality solo singing videos. We select 400 single-person English singing clips with unoccluded frontal faces for the Singing task. 3.2.3 Data Processing and Statistics. To accommodate the typical 10-second output window of current models, we standardized all clips to a uniform duration via precise trimming or padding. We retain only videos with a resolution of at least 720P to ensure high-quality visual input. In addition, all clips are stripped of their original audio so that models must rely solely on visual cues. After this processing, the final benchmark comprises 1,634 high- quality video clips. As shown in Figure 2, the SFX subset encom- passes 225 distinct sound events, spanning 10 major categories (Animals, Transport, Human Vocal, Sports, Household, Nature, In- dustrial, Alarms, Daily Activity, and Others) and 29 subcategories. The Instrument Performance subset includes 55 different instru- ment types, grouped into six categories: Strings, Winds, Percussion, Drums, Keyboards, and Electronic (Figure 3(a)). For the Speech and Singing subsets, we further analyze the apparent age and gender 3 distributions of the visible subjects, as shown in Figure 3(b). Further details are available in Appendix A.1. Household Transport Alarms Animals Industrial Sports Others Daily Nature Human Elec-devices Home appliances Doors/windows Road vehicles Vehicle signals Rail transport Air transport Water transport Weapons Domestic alerts Emergency sirens Mammals Birds Reptiles Insects Tools Processing Agricultural Ball sports Recreational activities Resonance Beat boxing Food-related Object interaction Water/ice Weather/wind Fire/geological Human vocal Human contact VidAudio-Bench SFX subset 29 subcategories 225 sound events Figure 2: Data distribution of the SFX subset. The inner ring denotes high-level categories, and the outer bars show the number of sound events in each subcategory. 3.3 V2A and VT2A Paradigms A primary challenge in V2A generation is the one-to-many map- ping between visual and acoustic signals. For instance, a beach scene could be paired with either ambient wave sounds or relaxing background music. This ambiguity complicates dimension-specific evaluations. To address this, we formalize two input configurations to ensure a controlled and category-specific assessment. V2A Setting (Task-Prompted): To eliminate categorical ambi- guity while strictly relying on visual cues for content generation, our baseline V2A setting employs a minimal task-level instruc- tion (e.g., "Realistic foley sound synchronized with the video"). This instruction serves merely as a categorical control signal, guiding the model toward the intended audio type without introducing ex- plicit semantic descriptions of the visual events. We assess whether this minimal instruction is correctly followed through a dedicated Instruction-Following dimension (see Section 4.3). VT2A Setting (Caption-Augmented): To investigate whether generation models truly understand the video content, we extend the benchmark to a Video-Text-to-Audio (VT2A) setting. In this setup, the input comprises the video and a prompt synthesized from visual content and task instructions. To prevent any potential information leakage from pre-existing audio-related text, we utilize Qwen3-VL [2] to extract descriptions strictly from muted videos. This forces the Vision-Language Model (VLM) to describe only the visual elements (e.g., actions, objects, and environment). These raw visual descriptions are subsequently formatted to match specific audio generation templates, providing only video-observable se- mantics. This approach enables us to explore the modelâs generative ability when provided with visual text, and to assess whether this (a)(b) Figure 3: (a) Distribution of categories in the Instrument Performance subset. (b) Gender and age group distributions in the Speech and Singing subsets. leads to meaningful improvements or provides semantic assistance, in comparison to the V2A setting. We assess the fidelity of the generated visual descriptions using a hybrid validation strategy. For tasks with original labels (SFX and Instrument Performance), LLM-based similarity evaluation yields a semantic retention accuracy of 70.4% under a 0.5 threshold. For tasks without discrete labels (BGM, Speech, and Singing), human evaluation on a 15% sample shows strong alignment overall, with scores of 0.83 for Speech, 0.81 for BGM, and 0.67 for Singing. These results confirm that our prompts provide effective visual semantics for VT2A generation. Detailed implementation of these two settings is provided in Appendix A.2. 4 Evaluation Metrics To comprehensively assess V2A generation, we propose a unified evaluation framework based on three complementary perspectives: (1) Audio Quality (AQ), which focuses on the intrinsic proper- ties of the generated audio, including fidelity, perceptual quality, and task-specific characteristics (e.g., musicality). (2) Video-Audio Consistency (VAC), which measures the alignment between au- dio and visual content in terms of semantic correspondence and temporal synchronization, as well as task-specific attributes such as affective alignment. (3) Text-Audio Consistency (TAC), which assesses whether the generated audio conforms to the intended instruction or expected semantic content. These three perspectives are further refined into thirteen fine-grained dimensions, tailored to the characteristics of each task, as shown in Figure 4. 4.1 Audio Quality Evaluation AQ - Fidelity. To evaluate audio fidelity, signal integrity, and sus- ceptibility to perceptual artifacts, we compute the FrĂŠchet Distance (FD) on Audio-MAE [19] embeddings. Compared to commonly used embeddings such as VGGish [37], Audio-MAE demonstrates su- perior Precision sensitivity, providing an empirical upper bound for detecting additive noise and filtering artifacts [25]. The refer- ence distribution is constructed using category-matched real audio datasets. Specifically, for SFX we use the sound event subset of VGG- Sound [6], while for Speech we use the AVSpeech [15] training set. This is necessary because FD measures distributional differences in the embedding space, and audio embeddings are highly dependent on semantic category and acoustic structure. Using mismatched ref- erence distributions would cause FD to reflect content distribution 4 Fidelity Aesthetic Temporal Sync Semantic Correspondence Semantic Alignment Instruction Following Fidelity Aesthetic Musicality Rhythmic Sync Temporal Sync Semantic Correspondence Instruction Following Semantic Alignment Intelligibility Perception Lip Sync Fidelity Semantic Correspondence Identity Consistency Instruction Following Semantic Alignment Affective Alignment Intelligibility Musicality Perception Semantic Correspondence Lip Sync Affective Alignment Semantic Alignment Fidelity Instruction Following Audio Quality Video-Audio Consistency Text-Audio Consistency SFX Music Speech Singing User Determine whether the visible person in the silent video and the human voice in the audio are demographically consistent in terms of apparent gender presentation and apparent age group. Step 1: Visual Analysis (Video Only) Estimate: 1....... Step 2: Acoustic Analysis (Audio Only) Estimate:2....... Step 3: Consistency Rules : a)Gender consistency: ...... b)Age consistency: ...... Step 4: Final Scoring : Score rules and evidence...... MLLM visual_profile: age_group: Child (0â12), gender_presentation: female- presenting, age_confidence: 0.9, gender_confidence: 0.9 , evidence: "The person is a young girl with a childlike facial structure. She is wearing a white sweatshirt and has her hair in a ponytail, which is common for a child.". vocal_profile: age_group : ...... MLLM-as-a-Judge Identity Consistency Figure 4: Overview of the evaluation framework in VidAudio-Bench. For each audio generation task, we define a set of evaluation dimensions across Audio quality, Video-Audio Consistency, and Text-Audio Consistency. An MLLM-as-a-Judge framework is employed to assess the generated audio through multi-step reasoning based on the input video and audio. differences rather than signal fidelity. Therefore, category-specific reference distributions allow FD to more accurately measure signal- level degradations while minimizing semantic distribution bias. AQ - Aesthetic. We utilize the Audiobox-Aesthetics [47] frame- work to assess the aesthetic quality of generated sound effects and music. This framework decomposes audio aesthetics into four key dimensions: Production Quality (PQ), Production Complex- ity (PC), Content Enjoyment (CE), and Content Usefulness (CU). For SFX and Instrumental Performance tasks, we prioritize PQ, CE, and CU as the primary metrics. PC is excluded in these cases be- cause the visual context inherently constrains the audioâs degrees of freedom, rendering complexity a less meaningful indicator. In contrast, for BGM, all four dimensions are considered. The weight- ing schemesâ4:3:3 (CE:PQ:CU) for SFX/Instrument Performance and 4:2:2:2 (CE:PQ:CU:PC) for BGMâare based on utterance-level Pear- son correlations in prior work [47], which we apply here for the first time to assign evaluation weights. AQ - Intelligibility. Intelligibility is essential for evaluating Speech and Singing. In this work, we use STOI-Net [65] to assess intelligibil- ity in a non-intrusive manner. While STOI-Net has been previously applied to speech, we employ it here for the first time in singing generation. Traditional intrusive metrics such as STOI [42], which require access to the original clean waveform, are unsuitable for generative tasks [64], as even a perfect model may not reproduce the original audio. Similarly, metrics like Character Error Rate (CER) and Word Error Rate (WER), which rely on reference transcriptions, are inapplicable for non-intrusive evaluation. AQ - Musicality. Following prior work [44], we evaluate musical- ity using three objective metrics: Pitch Class Histogram Entropy (PCE) [56] to quantify tonal clarity (where lower entropy indi- cates a more salient harmonic center), Grooving Pattern Similarity (GS) [56] to measure rhythmic regularity by comparing pattern consistency across bars, and Empty Beat Rate (EBR) [14] to assess note density by calculating the proportion of silent beats. Since V2A models generate raw audio, we first convert the outputs to MIDI via Basic Pitch [4]. We define a Validity Rate (í rate ) to account for samples yielding valid musical content. We formulate Musicality Score (MS) in Eq.(1)by normalizing all metrics to[0,1], where 1â PCE/log 2 12 specifically quantifies tonal clarity relative to a uniform 12-pitch distribution. MS=í rate ¡ GS+ 1â PCE log 2 12 +(1â EBR) 3 .(1) AQ - Perception. To assess perceptual speech quality in terms of naturalness, clarity, and overall listening experience, we em- ploy DNSMOS Pro [11], a non-intrusive probabilistic model for Mean Opinion Score (MOS) estimation. It adopts a lightweight end-to-end architecture to model the MOS posterior distribution, achieving high accuracy with reduced computational cost. We use SingMOS-Pro [43] to evaluate the perceptual quality and acoustic pleasantness of the generated singing voices. SingMOS-Pro is de- signed for automatic singing quality assessment, providing reliable MOS annotations of overall perceptual quality of singing vocals. 4.2 Video-Audio Consistency Evaluation VAC - Temporal Sync. We adopt the DeSync score predicted by the Synchformer [24] to quantify event-level audioâvideo synchro- nization, where lower absolute values indicate better alignment. Following MMAudio [9], we calculate the average offset using the first and last 4.8s of each video, allowing for overlap. This method evaluates both SFX and Instrument Performance tasks. VAC - Lip Sync. To evaluate audio-visual synchronization in Speech and Singing, we adopt LatentSync [28], which is designed for fine- grained lip-sync detection. We report the average absolute value of the offset, which measures temporal misalignment, where lower values indicate better synchronization. VAC - Rhythmic Sync. Standard tools like Synchformer are ill- suited for evaluating background music. Thus, we introduce a rhyth- mic synchronization score to jointly evaluate BGM rhythm similar- ity and temporal alignment. The formula is í rhythm = í + 1 2 ¡ exp(âíź|ÎíĄ|),(2) whereíis the Pearson correlation coefficient between the video motion envelope and the audio energy envelope, andÎíĄis the optimal temporal offset estimated by cross-correlation. We mapí to[0,1]before combining it with the temporal penalty term. The exponential factor penalizes large temporal offsets, whereíź= ln 2 í , 5 andírepresents the time threshold at which the penalty weight halves. Based on standard music theory and previous BGM genera- tion practices [13], we setí=0.5 seconds, corresponding to one full beat at a tempo of 120 BPM. Such misalignment indicates the visual motion and musical rhythm are out of sync. This formulation ensures that high synchronization scores require both strong corre- lation and minimal temporal misalignment, making it practical for reference-free evaluation of audio-visual rhythmic consistency. VAC - Semantic Correspondence. To evaluate the semantic con- sistency between visual and audio content, we employíźííĄíííííż â íźíľ + +(ííí.)from FreeBind [55]. By incorporating audio information from CLAP and fine-tuning the audio encoder, this space achieves improved audioâimage alignment and strong performance across audio tasks. It provides a reliable metric for assessing audioâvisual semantic consistency, surpassing the original ImageBind [17]. VAC - Identity Consistency. For Speech and Singing tasks, we adopt an MLLM-as-a-Judge framework (Figure 4), using Qwen3- Omni [60] as the judge model, to evaluate whether the generated voice is consistent with the visible person in the video. The evalua- tion focuses on demographic consistency, including apparent age and gender. Based on previous work [3], we categorize age groups into Child (0â12), Teenage (13â17), Adult (18â59), and Senior (60+). The detailed evaluation prompts are provided in Appendix B.1. VAC - Affective Alignment. Similarly, using the MLLM-as-a- Judge approach, we evaluate whether the emotion expressed in the generated speech or singing matches the emotion suggested by the visual scene. The detailed prompts are provided in Appendix B.2. 4.3 Text-Audio Consistency Evaluation TAC - Semantic Alignment. To evaluate the semantic consistency between textual content and generated audio, we adopt CLAP [57] to compute the cosine similarity of their embeddings. We use differ- ent checkpoints for different tasks to enhance the semantic align- ment between audio and text representations. TAC - Instruction Following. As defined in Section 3.3, each generation task is guided by a specific category instruction. It is crucial to assess whether the generation model faithfully follows the given instruction and produces audio of the intended category. To this end, we adopt the MLLM-as-a-Judge framework to systemat- ically verify instruction compliance. More detailed implementation information of the judge model can be found in Appendix B.3. 5 Experiments 5.1 Evaluated Audio Generation Models We evaluate 11 representative models on VidAudio-Bench, includ- ing 8 video-to-audio models and 3 video-to-music (V2M) models, covering 10 open-source models and 1 commercial model. The evaluated models are briefly summarized as follows: â˘AudioX [45] is a unified Diffusion Transformer (DiT) for anything- to-audio generation that supports diverse multimodal conditions, including video, text, and images. â˘FoleyCrafter [68] adapts a pre-trained T2A model for V2A generation with a semantic adapter and a temporal controller. â˘HunyuanVideo-Foley [39] is an end-to-end VT2A DiT that leverages self-supervised audio features and dual-stream fusion for high-fidelity synchronization. â˘Kling-Foley [51] is a DiT-based V2A model enhancing visual- semantic and temporal alignment for high-fidelity synthesis. ⢠MMAudio [9] is a multimodal framework improving V2A syn- thesis by jointly learning from text-audio and audio-visual data. â˘ReWaS [26] is a VT2A method that uses video as structural control and text prompts as semantic guidance. ⢠ThinkSound [31] introduces Chain-of-Thought reasoning into V2A generation for stepwise audio synthesis and editing. ⢠UniFlow-Audio [61] is a unified flow-matching framework em- ploying a dual-fusion mechanism for omni-modal alignment. ⢠GVMGen [72] is a V2M model using hierarchical attention for spatial-temporal alignment in zero-shot music generation. â˘SONIQUE [67] is a customizable V2M model that uses LLMs to bridge unpaired data by converting visual descriptions into musical tags for diffusion-based generation. ⢠VidMuse [46] is a V2M framework with long-short-term model- ing capturing both local and global visual cues. 5.2 Main Results Table 1 presents the performance of all models on VidAudio-Bench. Overall Findings. Model performance varies substantially across tasks and evaluation dimensions. Although many models perform reasonably well on SFX and Music, Speech and Singing remain much more challenging, with clear drops in perceptual quality. One likely reason is that current mainstream V2A training data (e.g., VGG- Sound [6] and AudioSet [16]) are dominated by environmental sound events, providing limited coverage of highly structured and semantically complex human vocalizations. (a)(b) (c)(d) AudioXThinkSoundKling-FoleyFoleyCrafter HunyuanVideo-FoleyUniFlow-AudioMMAudio ReWaS T-A Semantic-Align T-A Semantic-Align Instruction Following Instruction Following Instruction Following Instruction Following Fidelity Fidelity Fidelity Fidelity Aesthetic Aesthetic V-A Semantic-Corr V-A Semantic-Corr V-A Semantic-Corr V-A Semantic-Corr Identity- Cons Identity- Cons Lip- sync Lip- sync Intelligibility Intelligibility Perception Perception Musicality Musicality Affective- Align Affective- Align Temp-sync Temp-sync T-A Semantic-Align T-A Semantic-Align Figure 5: Task-wise radar plots of eight representative mod- els across four audio generation tasks using V2A results: (a) sound effects, (b) music, (c) speech, and (d) singing. Each sub- plot summarizes model performance over the task-specific evaluation dimensions. 6 Table 1: Evaluation results on VidAudio-Bench across multiple tasks and dimensions. Results are presented in the format of V2A / VT2A, whereâ(â) indicates higher (lower) values represent better performance. Best results in each category are highlighted in bold. Music-I: Instrument Performance; Music-B: Background Music. Dimensions and TasksFidelityâAestheticâIntelligibilityâ ModelsSFXMusic-IMusic-BSpeechSingingSFXMusic-IMusic-BSpeechSinging AudioX [45]5.480/9.67510.533/7.67025.297/19.726 7.300/5.60713.787/6.2174.587/4.7856.081/6.1334.485/5.5030.684/0.6260.613/0.593 FoleyCrafter [68]11.847/16.417 17.537/24.443 20.100/37.038 19.362/28.316 22.894/27.5294.560/4.9515.821/5.8925.207/4.3690.552/0.5560.542/0.528 HunyuanVideo-Foley [39]8.205/12.24014.369/7.721 10.187/22.904 6.796/21.03058.529/17.475 5.080/5.0606.189/6.0006.659/5.3850.583/0.5720.557/0.564 Kling-Foley [51]9.090/9.52918.585/17.114 19.767/19.150 35.304/35.005 12.981/13.1184.855/4.9155.957/6.1766.035/6.0060.554/0.5530.552/0.553 MMAudio [9] 13.499/12.347 17.952/16.763 24.060/31.994 14.402/13.279 20.109/16.0044.685/4.7406.161/6.0026.094/4.9370.548/0.5340.483/0.468 ReWaS [26]10.308/10.165 17.429/17.717 34.675/34.594 23.848/23.317 12.995/13.0074.510/4.5254.748/4.7284.353/4.3480.861/0.863 0.896/0.896 ThinkSound [31]4.943/4.335 6.582/6.56322.048/20.562 4.490/7.343 8.625/11.3704.653/4.6186.243/6.2826.047/4.7330.599/0.6000.572/0.577 UniFlow-Audio [61]16.214/17.484 17.004/27.911 48.409/54.746 69.896/68.009 51.483/48.7304.745/4.4095.534/4.4053.982/3.7270.528/0.5790.424/0.498 Dimensions and TasksV-A Semantic-CorrâTemp-SyncâRhy-SyncâLip-Syncâ ModelsSFXMusic-IMusic-BSpeechSingingSFXMusic-IMusic-BSpeechSinging AudioX [45]0.194/0.2290.238/0.2920.120/0.1840.157/0.2020.197/0.2151.268/1.2461.288/1.3430.123/0.1039.422/9.8099.660/9.115 FoleyCrafter [68]0.201/0.2260.205/0.2050.179/0.1900.164/0.2380.212/0.2411.248/1.2271.255/1.2970.154/0.1329.053/8.9169.040/9.184 HunyuanVideo-Foley [39]0.208/0.2240.296/0.2940.187/0.192 0.238/0.2430.228/0.2470.673/0.6060.375/0.3810.134/0.1742.351/2.2841.024/0.888 Kling-Foley [51] 0.234/0.2460.252/0.2830.139/0.1490.216/0.2180.257/0.2590.526/0.5300.394/0.3710.139/0.1632.062/2.0930.853/0.824 MMAudio [9]0.191/0.2190.302/0.3160.159/0.1510.220/0.2230.255/0.2320.480/0.457 0.260/0.265 0.172/0.1871.390/1.4640.733/0.818 ReWaS [26]0.072/0.0720.059/0.0600.076/0.0750.047/0.0480.063/0.0631.022/1.0361.154/1.1580.137/0.1387.556/7.8097.773/7.746 ThinkSound [31]0.199/0.1990.274/0.2890.150/0.1560.231/0.2320.232/0.2410.620/0.6290.382/0.3850.163/0.192 1.142/1.1510.968/0.861 UniFlow-Audio [61]0.212/0.1660.237/0.1710.153/0.1210.145/0.1220.139/0.1101.127/1.1851.267/1.2210.148/0.1399.600/9.6809.131/9.083 Dimensions and TasksT-A Semantic-AlignâInstruction-Followingâ ModelsSFXMusic-IMusic-BSpeechSingingSFXMusic-IMusic-BSpeechSinging AudioX [45]0.032/0.3220.231/0.4440.235/0.3500.341/0.4370.534/0.4140.795/0.8320.801/0.8430.294/0.5370.825/0.9810.838/0.878 FoleyCrafter [68]0.028/0.2800.288/0.3470.293/0.3320.323/0.3570.549/0.365 0.973/0.9650.921/0.8120.498/0.1950.316/0.9980.903/0.580 HunyuanVideo-Foley [39]0.001/0.3290.343/0.4470.247/0.375 0.404/0.3290.455/0.3620.475/0.988 0.953/0.8170.723/0.5190.993/0.9930.930/0.525 Kling-Foley [51]0.082/0.320 0.317/0.4560.282/0.2810.369/0.4000.478/0.4160.935/0.9420.753/0.8470.530/0.7040.959/0.9760.828/0.793 MMAudio [9]0.130/0.3520.263/0.4200.340/0.3130.269/0.3450.496/0.3380.935/0.9780.895/0.8370.628/0.3550.995/0.9980.935/0.553 ReWaS [26]0.001/0.0080.274/0.1410.173/0.1920.278/0.0400.159/0.1520.662/0.6580.440/0.4500.398/0.3900.209/0.2090.000/0.003 ThinkSound [31]0.066/0.2130.246/0.3820.305/0.2150.320/0.3610.478/0.2710.848/0.818 0.911/0.916 0.736/0.3160.956/0.9320.960/0.835 UniFlow-Audio [61]0.019/0.1520.335/0.3150.119/0.1950.348/0.2850.177/0.2630.920/0.9280.880/0.7230.052/0.0390.711/0.5150.050/0.040 Dimensions and TasksMusicalityâPerceptionâIdentity-ConsâAffective-Alignâ ModelsMusic-IMusic-BSingingSpeechSingingSpeechSingingSpeechSingingAvg.Rank AudioX [45]0.637/0.7140.700/0.7160.723/0.7412.072/2.2172.762/2.8534.388/4.6633.405/3.8252.187/2.3453.233/3.2035.462/3.769 FoleyCrafter [68]0.713/0.6350.646/0.5650.763/0.6742.271/2.8692.091/2.3653.325/4.6073.943/3.1382.403/3.0343.448/3.3434.795/5.179 HunyuanVideo-Foley [39] 0.751/0.652 0.751/0.6680.735/0.6833.037/2.9243.212/3.3074.901/4.9083.798/4.3252.420/2.4373.748/3.7953.128/3.359 Kling-Foley [51]0.614/0.6690.715/0.7300.694/0.7052.588/2.5913.513/3.5074.818/4.8233.928/3.9252.762/2.6973.943/3.8853.718/3.128 MMAudio [9]0.678/0.6560.703/0.6160.748/0.6422.983/2.874 3.625/3.3654.784/4.9444.128/4.5452.204/2.4883.808/3.6903.436/3.769 ReWaS [26]0.545/0.5510.579/0.5780.605/0.6062.048/2.0461.458/1.4603.782/3.8423.348/3.2652.354/2.3182.920/2.9436.513/6.590 ThinkSound [31] 0.716/0.7260.748/0.6780.746/0.7102.689/2.5993.005/3.0394.626/4.6143.170/3.4752.833/2.6483.873/4.0853.051/3.513 UniFlow-Audio [61]0.704/0.6940.542/0.5150.673/0.6511.721/1.6791.993/2.0054.473/4.4784.205/4.2632.726/2.4084.175/3.9755.820/6.641 We also observe that no single model consistently ranks best across all perspectives and dimensions. Instead, different models ex- hibit different strengths, suggesting inherent trade-offs among these objectives. For example, models that align better with visual con- tent (e.g., Kling-Foley) do not necessarily generate higher-quality audio, while models with stronger perceptual quality (e.g., AudioX) often fall short on fine-grained temporal alignment, such as lip synchronization. These results suggest that V2A/VT2A evaluation cannot be adequately captured by a single overall score, and instead requires a multi-dimensional, task-aware evaluation framework. Task-wise Results. We conduct detailed analyses across more than four task categories, as illustrated in Figure 5. Sound Effects. As shown in Figure 5(a), ThinkSound and AudioX lead in fidelity, MMAudio excels in temporal synchronization and textâaudio consistency, while Kling achieves stronger videoâaudio alignment. All the models demonstrate relatively good performance. Music. Figure 5(b) shows that for the Instrument Performance task, ThinkSound leads in fidelity and aesthetic quality, with melody performance second only to Hunyuan. MMAudio excels in tempo- ral synchronization and videoâaudio alignment, while Hunyuan achieves the best textâaudio consistency. Most models also success- fully follow the instructions to generate the intended music. For Table 2: Evaluation results of three specialized BGM genera- tion models on the BGM task across five dimensions. ModelsFidelityâ Aestheticâ Musicalityâ Semantic-Corrâ Rhy-Syncâ GVMGen [72]12.1716.5980.7460.1000.151 SONIQUE [67]34.8075.7670.6380.1130.093 VidMuse [46] 10.9607.0800.7170.1410.128 the BGM task, since most general-purpose models struggle with background music generation, we additionally compare three spe- cialized BGM generation models (Table 2). Hunyuan achieves the highest fidelity, followed by GVMGen and VidMuse, and exhibits stronger melodic structure than GVMGen. In terms of rhythmic synchronization, Kling performs the best. However, these special- ized models score lower on semantic consistency, indicating that optimizing for musical quality alone does not guarantee strong semantic alignment with visual or textual conditions. Speech. As shown in Figure 5(c), several models (e.g., ThinkSound) achieve strong performance in lip synchronization, fidelity, and feature consistency. Although ReWaS achieves the highest intelligi- bility, its remarkably low Instruction-Following score indicates a critical qualitative flaw. Specifically, while the model produces per- fectly articulated syllables, they form nonsensical gibberish when 7 UniFlow-AudioHunyuanVideo-FoleyKling-FoleyMMAudio VidAudio-Bench Human Figure 6: Human preference correlation with VidAudio-Bench. This figure shows the Pearson correlation coefficients (í) between VidAudio-Bench scores (x-axis) and human win rates (y-axis) across various evaluation dimensions. High correlations demonstrate strong alignment with human perceptual judgments. strung together, leading the MLLM to reject the audio as natural hu- man speech. These results reveal a clear trilemma in current speech generation models: jointly optimizing clarity, synchronization, and semantic alignment remains challenging. Singing. As shown in Figure 5(d), MMAudio achieves strong over- all vocal perceptual quality, with melody performance second only to FoleyCrafter, and also excels in synchronization but lags behind most models in intelligibility. Kling demonstrates stronger overall semantic consistency. Notably, no existing model can simultane- ously balance melodicity, lyric intelligibility, visual synchronization, and semantic alignment in singing tasks. V2A vs. VT2A. To investigate the effect of explicit visual descrip- tions on generation, we jointly analyze video-audio semantic cor- respondence (V-A Semantic-Corr) and instruction following (IF), where IF measures whether the generated audio matches the target category. As shown in Table 1, rather than improving, IF frequently drops when moving from the V2A setting to VT2A. This drop is particularly evident in complex audio categories such as BGM and Singing. For instance, MMAudio drops from 0.935 to 0.553 on Singing, and HunyuanVideo-Foley drops from 0.930 to 0.525. We attribute this degradation to two factors. First, longer visual descriptions can dilute the core instruction, making it harder for the model to preserve the target category when processing dense contextual details. As a result, the model may be distracted by sec- ondary information, such as objects, attributes, or scene elements, instead of following the main task requirement. Second, explicit visual descriptions may introduce semantic cues that conflict with the intended audio category. For instance, in the BGM task, descrip- tions of visible actions such as "a car speeding" can bias the model toward generating event-driven sound effects rather than back- ground music, leading to severe IF drops, such as UniFlow-Audio falling to 0.039 on BGM. Overall, these results point to a fundamen- tal tension in current V2A systems between category-level control and visually grounded generation. This also explains why some models achieve higher V-A semantic consistency in VT2A despite lower IF: even when they miss the intended category, the generated audio can still align closely with the visible content of the video. Table 3: Binary classification performance of the automated Instruction Following metric against human-labeled audio categories across four task types. Category Accuracyâ Precisionâ Recallâ F1-scoreâ SFX0.86250.85910.99220.9209 Music0.80000.72530.90410.8042 Speech0.88750.88550.97480.9280 Singing0.84380.80000.92680.8588 5.3 Human Evaluation Correlation Analysis In this section, we conduct large-scale human evaluations and com- pute the correlation between scores from automatic metrics and hu- mans to validate VidAudio-Benchâs alignment with human senses. Human Evaluation. For each task, we selected 20 representative videos. We then paired these with audio generated by 4 different models, resulting in a total of 400 audioâvideo pairs. To mitigate the influence of individual subjective preferences, each pair was evaluated by five independent raters. To avoid cross-task interfer- ence, each rater was assigned to evaluate only one specific type of audio task. In total, 20 participants were involved in the study. For each audio category, we focus on four key dimensions: semantics, synchronization, realism, and instruction following. After watching each video, raters scored from 1 to 5 for each of these dimensions. Evaluation Methodology. We employ three strategies to assess alignment with human preferences: (1) Pairwise Correlation: For metrics allowing pairwise comparison, we calculate win rates, as- signing 1 for a win, 0 for a loss, and 0.5 for a tie, for both human and benchmark scores, then compute the Pearson correlation (í) between them. (2) Direct Correlation: For metrics like Fidelity, we di- rectly correlate raw model scores with human ratings. As shown in Figure 6, these high correlations demonstrate strong alignment with human judgments. (3) Classification Accuracy: For the IF metric, we adopted a binary classification evaluation to assess the agreement between human-labeled categories and the MLLMâs predictions. The results are summarized in Table 3. High accuracy and F1-scores across all categories confirm that our automated IF metric closely mirrors human judgment in verifying instruction adherence. 8 6 Conclusion In conclusion, VidAudio-Bench establishes a comprehensive bench- mark for V2A/VT2A evaluation, built on 1,634 carefully curated video-text pairs with strong audio-visual correlation. Covering four audio types and thirteen evaluation dimensions, it supports reliable and interpretable assessment through automated, multidimensional, and human-aligned evaluation. The benchmark also reveals the central challenge of balancing Audio Quality, Video-Audio Consis- tency, and Text-Audio Consistency. VidAudio-Bench offers valuable insights for achieving more coherent and perceptually grounded audio generation, constituting a significant and robust contribution to research and evaluation in this field. References [1]Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. 2021. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in neural information processing systems 34 (2021), 24206â24221. [2]Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al.2025. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631 (2025). [3] Abigail Berthe-Pardo, Gaspard Michel, Elena V Epure, and Christophe Cerisara. 2026. S-VoCAL: A Dataset and Evaluation Framework for Inferring Speaking Voice Character Attributes in Literature. arXiv preprint arXiv:2603.00958 (2026). [4]Rachel M Bittner, Juan JosĂŠ Bosch, David Rubinstein, Gabriel Meseguer-Brocal, and Sebastian Ewert. 2022. A lightweight instrument-agnostic model for poly- phonic note transcription and multipitch estimation. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 781â785. [5]Zhe Cao, Tao Wang, Jiaming Wang, Yanghai Wang, Yuanxing Zhang, Jialu Chen, Miao Deng, Jiahao Wang, Yubin Guo, Chenxi Liao, et al.2025. T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation. arXiv preprint arXiv:2512.21094 (2025). [6]Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman. 2020. Vg- gsound: A large-scale audio-visual dataset. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 721â725. [7]Ziyang Chen, Daniel Geng, and Andrew Owens. 2024. Images that sound: Composing images and sounds on a single canvas. Advances in Neural Information Processing Systems 37 (2024), 85045â85073. [8] Ziyang Chen, Prem Seetharaman, Bryan Russell, Oriol Nieto, David Bourgin, Andrew Owens, and Justin Salamon. 2025. Video-guided foley sound generation with multimodal controls. In Proceedings of the Computer Vision and Pattern Recognition Conference. 18770â18781. [9]Ho Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya, Alexander Schwing, and Yuki Mitsufuji. 2025. Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis. In Proceedings of the Computer Vision and Pattern Recognition Conference. 28901â28911. [10] Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al.2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multi- modality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261 (2025). [11]Fredrik Cumlin, Xinyu Liang, Victor Ungureanu, Chandan KA Reddy, Christian SchĂźldt, and Saikat Chatterjee. 2024. DNSMOS Pro: A Reduced-Size DNN for Probabilistic MOS of Speech. In Interspeech. [12]Yusheng Dai, Zehua Chen, Yuxuan Jiang, Baolong Gao, Qiuhong Ke, Jun Zhu, and Jianfei Cai. 2026. Omni2Sound: Towards Unified Video-Text-to-Audio Generation. arXiv preprint arXiv:2601.02731 (2026). [13]Shangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang, Leyan Zhu, Zexin He, Hong- ming Liu, and Shuicheng Yan. 2021. Video background music generation with controllable music transformer. In Proceedings of the 29th ACM International Conference on Multimedia. 2037â2045. [14] Hao-Wen Dong, Wen-Yi Hsiao, and Yi-Hsuan Yang. 2018. Pypianoroll: Open source Python package for handling multitrack pianoroll. Proc. ISMIR. Late- breaking paper (2018). [15]Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein. 2018. Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation. ACM Transactions on Graphics (TOG) 37, 4 (2018), 1â11. [16]Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. 2017. Audio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 776â780. [17]Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. Imagebind: One embedding space to bind them all. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 15180â15190. [18]Daili Hua, Xizhi Wang, Bohan Zeng, Xinyi Huang, Hao Liang, Junbo Niu, Xinlong Chen, Quanqing Xu, and Wentao Zhang. 2025. Vabench: A comprehensive benchmark for audio-video generation. arXiv preprint arXiv:2512.09299 (2025). [19]Po-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, and Christoph Feichtenhofer. 2022. Masked autoencoders that listen. Advances in neural information processing systems 35 (2022), 28708â 28720. [20] Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren, Luping Liu, Mingze Li, Zhenhui Ye, Jinglin Liu, Xiang Yin, and Zhou Zhao. 2023. Make-an-audio: Text- to-audio generation with prompt-enhanced diffusion models. In International Conference on Machine Learning. PMLR, 13916â13932. [21]Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al.2024. Vbench: Comprehensive benchmark suite for video generative models. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 21807â21818. [22] Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al.2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024). [23]Vladimir Iashin and Esa Rahtu. 2021. Taming Visually Guided Sound Generation. In British Machine Vision Conference. BMVA Press. [24]Vladimir Iashin, Weidi Xie, Esa Rahtu, and Andrew Zisserman. 2024. Synch- former: Efficient synchronization from sparse cues. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 5325â5329. [25]Wonwoo Jeong. 2026. An Empirical Analysis of Task-Induced Encoder Bias in Fr\âechet Audio Distance. arXiv preprint arXiv:2602.23958 (2026). [26]Yujin Jeong, Yunji Kim, Sanghyuk Chun, and Jiyoung Lee. 2025. Read, watch and scream! sound generation from text and video. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 17590â17598. [27]Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre DĂŠfossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. [n. d.]. AudioGen: Textually Guided Audio Generation. In The Eleventh International Conference on Learning Representations. [28]Chunyu Li, Chao Zhang, Weikai Xu, Jingyu Lin, Jinghui Xie, Weiguo Feng, Bingyue Peng, Cunjian Chen, and Weiwei Xing. 2024. Latentsync: Taming audio- conditioned latent diffusion models for lip sync with syncnet supervision. arXiv preprint arXiv:2412.09262 (2024). [29] Susan Liang, Chao Huang, Filippos Bellos, Yolo Yunlong Tang, Qianxiang Shen, Jing Bi, Luchuan Song, Zeliang Zhang, Jason Corso, and Chenliang Xu. 2026. Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text- Conditioned Audio-Video Generation? arXiv preprint arXiv:2602.01623 (2026). [30]Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark Plumbley. 2023. AudioLDM: Text-to-Audio Generation with Latent Diffusion Models. In Proceedings of the 40th International Conference on Machine Learning, PMLR 2023, Vol. 202. International Machine Learning Society (IMLS), 21450â21474. [31]Huadai Liu, Kaicheng Luo, Jialei Wang, Wen Wang, Qian Chen, Zhou Zhao, and Wei Xue. [n. d.]. ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and Editing. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. [32]Yaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang, Yong Zhang, Haoxin Chen, Yang Liu, Tieyong Zeng, Raymond Chan, and Ying Shan. 2024. Evalcrafter: Benchmarking and evaluating large video generation models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 22139â22149. [33]Simian Luo, Chuanhao Yan, Chenxu Hu, and Hang Zhao. 2023. Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models. Advances in Neural Information Processing Systems 36 (2023), 48855â48876. [34]Jan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, and Soujanya Poria. 2024. Mustango: Toward controllable text-to- music generation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies (Volume 1: Long Papers). 8293â8316. [35] Juan F Montesinos, Venkatesh S Kadandale, and Gloria Haro. 2021. A cappella: Audio-visual singing voice separation. arXiv preprint arXiv:2104.09946 (2021). [36]Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PmLR, 8748â8763. [37]Dominik Roblek, Kevin Kilgour, Matt Sharifi, and Mauricio Zuluaga. 2019. Fr\âechet Audio Distance: A Reference-free Metric for Evaluating Music En- hancement Algorithms. In Proc. Interspeech. 2350â2354. 9 [38]Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016. Improved techniques for training gans. Advances in neural information processing systems 29 (2016). [39] Sizhe Shan, Qiulin Li, Yutao Cui, Miles Yang, Yuehai Wang, Qun Yang, Jin Zhou, and Zhao Zhong. 2025. Hunyuanvideo-foley: Multimodal diffusion with rep- resentation alignment for high-fidelity foley audio generation. arXiv preprint arXiv:2508.16930 (2025). [40]Roy Sheffer and Yossi Adi. 2023. I hear your true colors: Image guided audio gen- eration. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1â5. [41]Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, Yuxuan Wang, and Chao Zhang. 2024. video-SALMONN: speech-enhanced audio-visual large language models. In Proceedings of the 41st International Conference on Machine Learning. 47198â47217. [42]Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen. 2011. An algorithm for intelligibility prediction of timeâfrequency weighted noisy speech. IEEE Transactions on audio, speech, and language processing 19, 7 (2011), 2125â2136. [43]Yuxun Tang, Lan Liu, Wenhao Feng, Yiwen Zhao, Jionghao Han, Yifeng Yu, Jiatong Shi, and Qin Jin. 2025. SingMOS-Pro: An Comprehensive Benchmark for Singing Quality Assessment. arXiv preprint arXiv:2510.01812 (2025). [44] Sida Tian, Can Zhang, Wei Yuan, Wei Tan, and Wenjie Zhu. 2025. Xmusic: Towards a generalized and controllable symbolic music generation framework. IEEE Transactions on Multimedia 27 (2025), 6857â6871. [45]Zeyue Tian, Yizhu Jin, Zhaoyang Liu, Ruibin Yuan, Xu Tan, Qifeng Chen, Wei Xue, and Yike Guo. 2025. Audiox: Diffusion transformer for anything-to-audio generation. arXiv preprint arXiv:2503.10522 (2025). [46] Zeyue Tian, Zhaoyang Liu, Ruibin Yuan, Jiahao Pan, Qifeng Liu, Xu Tan, Qifeng Chen, Wei Xue, and Yike Guo. 2025. Vidmuse: A simple video-to-music genera- tion framework with long-short-term modeling. In Proceedings of the Computer Vision and Pattern Recognition Conference. 18782â18793. [47] Andros Tjandra, Yi-Chiao Wu, Baishan Guo, John Hoffman, Brian Ellis, Apoorv Vyas, Bowen Shi, Sanyuan Chen, Matt Le, Nick Zacharov, et al.2025. Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound. arXiv preprint arXiv:2502.05139 (2025). [48]Ilpo Viertola, Vladimir Iashin, and Esa Rahtu. 2025. Temporally aligned audio for video with autoregression. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1â5. [49] Hui Wang, Cheng Liu, Junyang Chen, Haoze Liu, Yuhang Jia, Shiwan Zhao, Jiaming Zhou, Haoqin Sun, Hui Bu, and Yong Qin. 2026. Tta-bench: A compre- hensive benchmark for evaluating text-to-audio models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 33512â33520. [50]Heng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright, and Weidong Cai. 2024. V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 15492â15501. [51]Jun Wang, Xijuan Zeng, Chunyu Qiang, Ruilong Chen, Shiyao Wang, Le Wang, Wangjing Zhou, Pengfei Cai, Jiahui Zhao, Nan Li, et al.2025. Kling-foley: Multi- modal diffusion transformer for high-quality video-to-audio generation. arXiv preprint arXiv:2506.19774 (2025). [52] Le Wang, Jun Wang, Chunyu Qiang, Feng Deng, Chen Zhang, Di Zhang, and Kun Gai. 2025. Audiogen-omni: A unified multimodal diffusion transformer for video- synchronized audio, speech, and song generation. arXiv preprint arXiv:2508.00733 (2025). [53]Yongqi Wang, Wenxiang Guo, Rongjie Huang, Jiawei Huang, Zehan Wang, Fuming You, Ruiqi Li, and Zhou Zhao. 2024. Frieren: Efficient video-to-audio generation network with rectified flow matching. Advances in neural information processing systems 37 (2024), 128118â128138. [54]Zehan Wang, Ke Lei, Chen Zhu, Jiawei Huang, Sashuai Zhou, Luping Liu, Xize Cheng, Shengpeng Ji, Zhenhui Ye, Tao Jin, et al.2025. T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. 23535â23547. [55]Zehan Wang, Ziang Zhang, Xize Cheng, Rongjie Huang, Luping Liu, Zhenhui Ye, Haifeng Huang, Yang Zhao, Tao Jin, Peng Gao, et al.2024. FreeBind: free lunch in unified multimodal space via knowledge fusion. In Proceedings of the 41st International Conference on Machine Learning. 52233â52246. [56]Shih-Lun Wu and Yi-Hsuan Yang. 2020. The jazz transformer on the front line: Exploring the shortcomings of ai-composed music through quantitative measures. arXiv preprint arXiv:2008.01307 (2020). [57]Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2023. Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1â5. [58]Zhifeng Xie, Shengye Yu, Qile He, and Mengtian Li. 2024. Sonicvisionlm: Playing sound with vision language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 26866â26875. [59]Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang, and Qifeng Chen. 2024. Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7151â7161. [60]Jin Xu, Zhifang Guo, Hangrui Hu, Yunfei Chu, Xiong Wang, Jinzheng He, Yuxuan Wang, Xian Shi, Ting He, Xinfa Zhu, et al.2025. Qwen3-omni technical report. arXiv preprint arXiv:2509.17765 (2025). [61]Xuenan Xu, Jiahao Mei, Zihao Zheng, Ye Tao, Zeyu Xie, Yaoyun Zhang, Haohe Liu, Yuning Wu, Ming Yan, Wen Wu, et al.2025. Uniflow-audio: Unified flow matching for audio generation from omni-modalities. arXiv preprint arXiv:2509.24391 (2025). [62]Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu. 2023. Diffsound: Discrete diffusion model for text-to-sound generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2023), 1720â1733. [63]Geng Yang, Shan Yang, Kai Liu, Peng Fang, Wei Chen, and Lei Xie. 2021. Multi- band melgan: Faster waveform generation for high-quality text-to-speech. In 2021 IEEE Spoken Language Technology Workshop (SLT). IEEE, 492â498. [64]Yochai Yemini, Aviv Shamsian, Lior Bracha, Sharon Gannot, and Ethan Fetaya. [n. d.]. LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading. In The Twelfth International Conference on Learning Representations. [65] Ryandhimas E Zezario, Szu-Wei Fu, Chiou-Shann Fuh, Yu Tsao, and Hsin-Min Wang. 2020. STOI-Net: A deep learning based non-intrusive speech intelligi- bility assessment model. In 2020 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC). IEEE, 482â486. [66] Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-llama: An instruction-tuned audio-visual language model for video understanding. In Proceedings of the 2023 conference on empirical methods in natural language processing: system demonstrations. 543â553. [67] Liqian Zhang and Magdalena Fuentes. 2025. Sonique: Video background music generation using unpaired audio-visual data. In ICASSP 2025-2025 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1â5. [68]Yiming Zhang, Yicheng Gu, Yanhong Zeng, Zhening Xing, Yuancheng Wang, Zhizheng Wu, Bin Liu, and Kai Chen. 2026. Foleycrafter: Bring silent videos to life with lifelike and synchronized sounds. International Journal of Computer Vision 134, 1 (2026), 46. [69]Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui, and Tamara L Berg. 2018. Visual to sound: Generating natural sound for videos in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition. 3550â3558. [70] Zitang Zhou, Ke Mei, Yu Lu, Tianyi Wang, and Fengyun Rao. 2025. Harmonyset: A comprehensive dataset for understanding video-music semantic alignment and temporal synchronization. In Proceedings of the Computer Vision and Pattern Recognition Conference. 3152â3162. [71]Alon Ziv, Itai Gat, Gael Le Lan, Tal Remez, Felix Kreuk, Jade Copet, Alexandre DĂŠfossez, Gabriel Synnaeve, and Yossi Adi. [n. d.]. Masked Audio Generation using a Single Non-Autoregressive Transformer. In The Twelfth International Conference on Learning Representations. [72] Heda Zuo, Weitao You, Junxian Wu, Shihong Ren, Pei Chen, Mingxu Zhou, Yujia Lu, and Lingyun Sun. 2025. Gvmgen: A general video-to-music generation model with hierarchical attentions. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 23099â23107. [73]Daniil Zverev, Thaddäus Wiedemer, Ameya Prabhu, Matthias Bethge, Wieland Brendel, and A Koepke. 2025. Vggsounder: Audio-visual evaluations for founda- tion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 1027â1037. 10 A VidAudio-Bench Construction To support reliable evaluation of V2A and VT2A systems, we con- struct subsets with clear audio-visual grounding, temporally com- plete events, and minimal confounding factors such as background music, narration, and static imagery. In this section, we describe the subset construction procedure of VidAudio-Bench and provide additional details of the V2A and VT2A settings. A.1 Subset Construction and Statistics SFX and Instrument Performance subsets. The SFX and Instru- ment Performance subsets are constructed from VGGSounder. We first enforce a strict audio-visual consistency criterion by re- taining only labels whose modality is annotated as AV, indicating that the sound source is visually observable in the video. We further exclude videos annotated withstatic_image,background_music, orvoice_over, thereby removing samples containing static im- agery, background music, or voice-overs. From the remaining can- didates, we retain only videos containing either a single dominant sound event or a clearly identifiable primary sound source. Based on semantic attributes, we group the 300 candidate labels into five categories and manually verify the grouping: sound effects, instrument, singing, speech, and others. Among them, 225 labels belong to sound effects, 55 to instrument, 6 to singing, 3 to speech, and 11 to others. For the SFX subset, we sample approximately 1â2 videos from each label in the sound effects category, resulting in 400 videos in total. For the Instrument Performance subset, we sample approx- imately 3â4 videos from each label in the instrument category, resulting in 191 videos in total. BGM subset. The BGM subset is constructed from the test split of HarmonySet. We first retain videos with durations in the range of 9â11 seconds, and then randomly sample 231 videos from the filtered set. To preserve semantic and temporal completeness, each selected video is further adjusted by slight padding or trimming to 10 seconds when necessary. The category distribution of the resulting subset is shown in Figure 7. 14.7% 20.3% 27.7% 23.8% 3.0% 10.4% Travel & Events (34) Life & Emotions (47) Arts & Performance (64) Technology & Fashion (55) Knowledge (7) Sports & Fitness (24) Figure 7: Content category distribution of the BGM subset. Speech subset. The Speech subset is constructed from the test split of AVSpeech. We first retain videos with durations of 9â11 seconds and standardize them to a uniform 10 seconds by padding or trimming. We then manually remove samples in which the speaker is visually too small for reliable lip-motion perception, as such videos usually contain unclear or barely visible lip movements. After this filtering process, 412 videos are retained in the final Speech subset. Singing subset. The Singing subset is constructed from the test split of Acappella. We first retain English singing videos and then segment the original videos into 10-second clips. From these clips, we randomly sample 400 videos to form the final subset. A.2 V2A and VT2A Settings In this section, we provide additional details on the prompt design used in our benchmark. Appendix A.2.1 describes the instruction templates for the two evaluation settings, while Appendix A.2.2 presents the prompts used for visual captioning with Qwen3-VL. A.2.1 Instruction Design for V2A and VT2A. Table 4 summarizes the instruction templates used in the V2A and VT2A settings. The prompting strategy is designed according to the capabilities of dif- ferent models. For models that support negative prompting (e.g., AudioX, HunyuanVideo, and MMAudio), we include negative in- structions to explicitly constrain the generation and better guide the output toward the target audio category. For models that do not support negative conditioning (e.g., FoleyCrafter, ReSound, ThinkSound, UniFlow-Audio, and Kling), we use only positive in- structions. This design choice is motivated by the observation that explicitly mentioning undesired audio concepts may unintention- ally bias the generation process, thereby increasing the risk of irrelevant or hallucinated audio content. In the VT2A setting, we use Qwen3-VL for visual caption generation, and Qwen3.5 only to rewrite the VT2A prompts into more natural text descriptions. Table 4: Predefined positive and negative instruction tem- plates used for the V2A and VT2A settings across four audio categories in VidAudio-Bench, where the Music category is further divided into Instrument Performance and BGM. DomainTaskPositive Instruction (In- put Text) Negative Instruction SFX V2A Realistic foley sound syn- chronized with the video. music, background music, speech, singing VT2ARealistic foley sound of cap- tion. music, background music, speech, singing Music- Instrument Performance V2A Musical instrument perfor- mance synchronized with the video. speech,singing,human voice VT2AInstrument performance of caption. speech,singing,human voice Music- BGM V2ABackground music matching the video scene. speech, sound effects, foley VT2ABackground music fitting the scene: caption. speech, sound effects, foley Speech V2AHuman speech synchro- nized with the video. background music, singing, noise, sound effects VT2A Speech by caption.background music, singing, noise, sound effects Singing V2A cappella singing voice syn- chronized with the video. speech, talking, instrumen- tal only, heavy accompani- ment VT2ASinging performance by caption. speech, talking, instrumen- tal only, heavy accompani- ment 11 A.2.2 Visual Captioning Prompt Design. To generate visual cap- tions that are better matched to the semantic characteristics of different audio categories, we adopt a category-specific prompting strategy for the Vision-Language Model (as shown in Figures 8â12). Specifically, different prompts are designed for sound effects, music, speech, and singing videos, with the music category further divided into instrument-performance and background-music scenarios. By explicitly adapting the instruction focus to each category, the model is encouraged to attend to the most relevant visual cues and avoid category-irrelevant or speculative descriptions, resulting in cap- tions that are more accurate, informative, and semantically aligned with the video content. Prompt Design for SFX Visual Captioning #Role : You are a precise visual event captioning system. Your task is to identify the dominant visible event that is responsible for sound production in the video. #Task: Describe the main visible sound-producing action in the video. #Constraints: - Only describe what is visually observable. - Do not infer off-screen events, intentions, or causes not supported by the video. - Do not describe acoustic attributes such as loudness, pitch, or timbre. - Be specific about the action, motion intensity, and environment when visible. - Use a highly concise telegraphic style: Subject + Action + Object/Environment. - Do not use introductory phrases such as âThe video shows,â âI can see,â or âA scene of.â - Output only one sentence describing the single dominant sound-producing event. #Example Output: A glass cup shattering on a tiled floor. A red sports car accelerating on an asphalt road. A person knocking on a wooden door. Figure 8: Prompt design for SFX visual captioning Prompt Design for Instrument Performance Visual Captioning #Role : You are a precise visual captioning system for instrument performance videos. Your task is to identify the main visible instrument performance event and describe it concisely. #Task: Generate one sentence describing the dominant instrument performance visible in the video. #Constraints: - Only describe information that is visually observable in the video. - Identify the instrument as specifically as possible only when supported by clear visual evidence; otherwise, use a more general instrument category. - Describe the visible playing action or technique when applicable. - Include the setting only when it is visually clear. - Do not infer the song title, genre, performer identity, brand, emotion, or any acoustic property not directly visible. - Format: [Subject] + [Action] + [Instrument] + [Setting]. #Example Output: A person strumming an electric guitar on a stage. A woman bowing a cello in a practice room. Hands pressing piano keys in a close-up indoor shot. Figure 9: Prompt design for Instrument Performance visual captioning Prompt Design for BGM Visual Captioning #Role : You are a precise visual captioning system. Your task is to objectively describe what is visually observable in the video. #Task: Describe the visual content of the video. #Constraints: - Describe only what is visually observable. - Include the main subjects, actions, and environment when clearly visible. - Describe the visual pacing or motion if visible (e.g., rapid movement, slow motion, steady camera). - Mention visible lighting, colors, or weather if they are clearly present. - Do not infer emotion, mood, atmosphere, intent, or story context beyond visible evidence. - Do not mention audio, music, sound effects, or any acoustic attribute. - Do not use introductory phrases such as âThe video shows,â âI can see,â or âA scene of.â - Output exactly one sentence. #Output: A cyclist riding along a tree-lined road in daylight. People dancing together on a brightly lit stage. A boat moving across a calm lake under cloudy skies. Figure 10: Prompt design for BGM visual captioning Prompt Design for Speech Visual Captioning #Role : You are a visual character profiler. Your task is to describe the dominant visible speaker and scene in the video. #Task: Describe the person who is speaking and the environment. #Constraints: - Describe gender, approximate age, facial expression, and emotional state. - Mention the environment or scene context only when it is visually evident. - Do not transcribe or guess the topic of what they are saying. - Do not guess the intent of their speech; focus solely on visual and environmental details. - Do not use introductory phrases such as âThe video shows,â âI can see,â or âA scene of.â - Output exactly one sentence. #Example Output: A man in his 50s, with a stern expression, speaking in a large conference room. A young woman with a bright smile, speaking outdoors on a sunny day. A middle-aged man with glasses, calmly speaking in a quiet library. Figure 11: Prompt design for Speech visual captioning Prompt Design for Singing Visual Captioning #Role : You are a precise visual captioning system. Your task is to describe the dominant visible singer and performance scene in the video. #Task: Describe the person singing in the video. #Constraints: - Only describe information that is visually observable. - Describe the singer or performers using directly supported cues such as clothing, pose, mouth movement, body movement, or microphone use. - Mention the setting or scene context only when it is visually evident. - Do not transcribe, paraphrase, or infer lyrics, song title, genre, identity, or intent. - Do not use introductory phrases such as âThe video shows,â âI can see,â or âA scene of.â - Output exactly one sentence. #Example Output: A singer holding a microphone and moving energetically on a brightly lit stage. A singer swaying gently while performing outdoors. A choir standing in formation and singing on a church stage. Figure 12: Prompt design for Singing visual captioning 12 B Evaluation Metrics In this section, we present the prompts used for the three evaluation methods involving MLLM-as-a-Judge discussed in the main text. B.1 Identity Consistency To evaluate cross-modal Identity Consistency, we adopt a structured protocol that examines whether the visible person in the video and the human voice in the audio are demographically aligned. Specifically, the evaluation is conducted from two complementary aspects, i.e., apparent gender presentation and apparent age group. The framework first analyzes visual cues from the silent video and acoustic cues from the audio independently to construct the visual and vocal demographic profiles. It then compares the two profiles according to predefined consistency rules and assigns a final score based on the degree of agreement between modalities. This design enables a systematic assessment of whether the gener- ated voice matches the on-screen person at the demographic level, while reducing interference from other factors such as semantic content, emotion, or audio quality. The detailed prompt used for this evaluation is illustrated in Figure 13. B.2 Affective Alignment Following the same evaluation framework as above, we further assess cross-modal Affective Alignment, focusing on whether the emotion expressed in the silent video is consistent with that con- veyed by the audio. In this setting, the evaluation considers two affective dimensions, namely emotion category and emotional in- tensity. Different emotion label sets are adopted for Speech and Singing. For speech, we use seven emotion categories, namely calm, happy, sad, angry, fearful, surprised, and disgusted, as shown in Figure 14. For singing, we use five categories, namely calm, happy, sad, angry, and fearful, as shown in Figure 15. B.3 Instruction Following For instruction-following evaluation, we focus exclusively on category- level compliance of the generated audio. Specifically, for each gen- erated sample, we assess only whether it belongs to the target audio category specified by the instruction, namely Speech, Singing, Music, and Sound Effects (SFX). This evaluation aims to verify whether the model produces audio of the intended category, without consider- ing finer-grained semantic correctness or perceptual quality. The detailed prompt used for this evaluation is illustrated in Figure 16. Prompt Design for Identity Consistency #Role : You are a professional multimodal evaluator for audiovisual demographic consistency analysis. Follow the evaluation protocol strictly. Use only the predefined labels. Output valid JSON only. Do not add extra text. Base all judgments on observable visual and acoustic cues. #Task: Determine whether the visible person in the silent video and the human voice in the audio are demographically consistent in terms of apparent gender presentation and apparent age group. The audio may contain either spoken voice or singing voice. In both cases, estimate demographic traits from the human vocal characteristics only. # Step 1: Visual Analysis (Video Only) Estimate: 1. apparent gender presentation [male-presenting / female-presenting / ambiguous] 2. apparent age group [Child (0â12), Teenage (13â17), Adult (18â59), Senior (60+)] 3. gender confidence [0.0-1.0] 4. age confidence [0.0-1.0] Use only visual cues such as face structure, skin appearance, hair, lip movement, overall visible age impression, and other observable appearance cues. # Step 2: Acoustic Analysis (Audio Only) Estimate: 1. apparent gender presentation [male-presenting / female-presenting / ambiguous] 2. apparent age group [Child (0â12), Teenage (13â17), Adult (18â59), Senior (60+)] 3. gender confidence [0.0-1.0] 4. age confidence [0.0-1.0] Use only acoustic cues such as pitch / F0 impression, timbre, resonance, vocal stability, vocal maturity, and overall voice-based age impression. Ignore semantic content, emotion, and audio quality. Focus only on demographic cues in the voice. # Step 3: Consistency Rules 1. Gender consistency: - same label -> "match" - different male-presenting / female-presenting labels -> "mismatch" - any ambiguous case -> "uncertain" 2. Age consistency: - same age group -> "match" - adjacent age groups -> "close" - distance >= 2 groups -> "mismatch" - weak or unclear evidence -> "uncertainâ # Step 4: Final Scoring - 5 = gender match + age match - 4 = gender match + age close - 3 = uncertain case due to weak, ambiguous, or low-confidence demographic evidence - 2 = reliable mismatch in either gender or age - 1 = reliable mismatch in both gender and age Uncertain case: - the sample is analyzable, but the demographic evidence is weak - gender is ambiguous - age is hard to judge reliably - both modality judgments exist but confidence is low # Output Requirement Return valid JSON only in exactly this format: âvisual_profileâ: âage_groupâ: ââ, âgender_presentationâ: ââ, ... , âvocal_profileâ: âage_groupâ: ââ, âgender_presentationâ: ââ, ... , "matching_analysis": "gender_consistency": "", "age_consistency": "", ... , "final_result": "final_score": 0 Figure 13: Prompt design for Identity Consistency. 13 Table 5: Evaluation dimensions and task-specific criteria used in the human subjective study across four task categories. Each dimension was rated on a 5-point Likert scale (1â5), and the table summarizes the aspect emphasized for each task. DimensionSFXMusicSpeechSinging RealismFidelity: Focuses on the clarity, fidelity, and absence of artifacts (e.g., noise, distortion) of the generated audio. SemanticsV-A Semantic-Corr: Match between audio and visual sound sources or key events. V-A Semantic-Corr: Match with specific instruments or overall visual atmosphere. Identity-Cons: Match be- tween the voice and the per- sonâs age and gender. Identity-Cons: Match be- tween the voice and the per- sonâs age and gender. SynchronizationTemp-Sync: Temporal align- ment between visual actions and sound onsets. Temp-Sync: Alignment with performance movements or rhythmic tempo. Lip-Sync: Synchronization between mouth motion and speech articulation. Lip-Sync: Synchronization between lip movements and melodic vocalization. Instruction FollowingEnvironmental/Event sounds; minimal music or speech. Melodic/Harmonic content; absence of distinct speech or dialogue. Intelligible human speech; no dominant music or back- ground SFX. Melodic vocal perfor- mance; distinct from plain speech or pure music. Prompt Design for Affective Alignment Category:Speech #Role : You are a professional multimodal evaluator for audiovisual emotion alignment. Follow the evaluation protocol strictly. Use only predefined labels. Output must be valid JSON only. No extra text. #Task: Evaluate whether the emotion in a silent video matches the emotion in the audio. # Step 1: Visual Analysis (Video Only) - Select emotion from: [calm, happy, sad, angry, fearful, surprised, disgusted] - Rate intensity (1-5): 1 very weak, 2 mild, 3 moderate, 4 strong, 5 extreme - Describe key motion: (face, lips, body) # Step 2: Acoustic Analysis (Audio Only) - Select emotion from the same set - Rate intensity (1-5) using the same scale - Describe key vocal features: (pitch, loudness, rhythm) # Step 3: Consistency Evaluation - consistent: same emotion AND similar intensity - partial: different emotion but same valence (e.g., sad vs angry) - inconsistent: opposite valence (e.g., happy vs sad) or extreme intensity gap - uncertain: unclear or low confidence # Step 4: Scoring (1-5) - 5: Perfect match in category and intensity. - 4: Same emotion, minor intensity difference. - 3: Different emotions but similar "vibe" (valence). - 2: Clear mismatch in emotion or intensity. - 1: Opposite emotions (e.g., laughing face with crying voice). # Output Requirement Return strictly in JSON: "visual_emotion": "", "visual_intensity": 0, "audio_emotion": "", "audio_ intensity ": 0, "consistency": "", "score": 0, "reason": "One sentence explanation" Figure 14: Prompt design for Affective Alignment on Speech. Prompt Design for Affective Alignment Category:Singing #Role : You are a professional multimodal evaluator for videoâsinging emotion alignment. Follow the evaluation protocol strictly. Use only predefined labels. Output must be valid JSON only. No extra text. #Task: Evaluate whether the emotion expressed in a silent video matches the emotion conveyed by the singing audio. # Step 1: Visual Analysis (Video Only) - Select emotion from: [calm, happy, sad, angry, fearful] - Rate intensity (1-5): 1 very weak, 2 mild, 3 moderate, 4 strong, 5 extreme - Describe key motion: (face, lips, body, gesture) # Step 2: Acoustic Analysis (Audio Only) - Select emotion from the same set: [calm, happy, sad, angry, fearful] - Rate intensity (1-5) using the same scale - Describe key singing cues: (melody, pitch contour, loudness, tempo, rhythm, timbre, vibrato, phrasing, articulation) - Focus on emotional expression in the singing voice and musical delivery, not audio quality - If lyrics are unclear, rely on vocal expression and musical cues only # Step 3: Consistency Evaluation - consistent: same emotion AND similar intensity - partial: different emotion but similar affective tone - inconsistent: opposite or clearly conflicting emotion, or extreme intensity gap - uncertain: emotion is unclear, ambiguous, or confidence is low # Step 4: Scoring (1-5) - 5: Perfect match in emotion category and intensity. - 4: Same emotion, minor intensity difference. - 3: Different emotions but similar affective feeling. - 2: Clear mismatch in emotion or intensity. - 1: Opposite or strongly conflicting emotions. # Output Requirement Return strictly in JSON: "visual_emotion": "", "visual_intensity": 0, "singing_emotion": "", "singing_intensity": 0, "consistency": "", "score": 0, "reason": "One sentence explanation" Figure 15: Prompt design for Affective Alignment on Singing. 14 Prompt Design for Instruction Following #Role : You are a deterministic audio classification model for audio content categorization. Follow the instructions strictly. Use only the predefined categories. Output must be valid JSON only. No extra text. If uncertain, choose the closest category. #Task: Classify the input audio into ONE dominant category: Speech / Singing / Music / Sound Effects (SFX) # Category Definitions 1. **Speech**: Human vocalizations conveying linguistic meaning, such as dialogue, monologue, narration, or whispering. - Key feature: phonetic structure and decodable language content. 2. **Singing**: Human vocalizations with sustained musical pitch, melodic contour, and song-like delivery, with or without lyrics. - Key feature: stable pitch, melody, and musical phrasing. 3. **Music**: Organized sound primarily produced by instruments or synthesized musical sources. - Key feature: melodic progression, harmonic structure, or intentional rhythmic composition. 4. **SFX**: Non-musical, non-linguistic sounds. Includes: - Environmental sounds: wind, rain, traffic, crowd murmur without distinct speech - Mechanical/physical sounds: engines, footsteps, impacts, glass breaking, ticking - Biological non-semantic sounds: coughing, sneezing, breathing, laughter, screams, animal calls # Conflict Resolution - Rhythmic but mechanical sounds (e.g., clock, metronome, machine loop) -> SFX - Vocal but non-semantic sounds (e.g., scream, laughter, cough, sobbing) -> SFX - Clear sustained melody and pitched vocal delivery dominate -> Singing - Clear linguistic content dominates without stable melody -> Speech - Humming with clear melody -> Singing - Instrumental-only audio -> Music # Internal Analysis Cues Before the final decision, consider internally: 1. Harmonicity: Is there stable pitch and overtone structure? 2. Temporal pattern: Is it impulsive, repetitive, continuous, or musically structured? 3. Semantic content: Can human language be decoded? 4. Source type: Is it vocal, instrumental, environmental, mechanical, or biological? 5. Intent: Does it sound composed/musical or natural/incidental? # Output Requirement Return valid JSON only in exactly this format: "category": "Speech" or "Singing" or "Music" or "SFX", "confidence": 0.00, "key_evidence": "brief reason for classification", "secondary_presence": "Speech" or "Singing" or "Music" or "SFX" or "None" Figure 16: Prompt design for Instruction Following. 15 Figure 17: Annotation interface used in the human subjective study. C Human Subjective Study We recruited a total of 20 participants with normal vision and hear- ing. To ensure high-quality feedback, the participants were divided into four groups, with 5 experts assigned to each task category (SFX, Music, Speech, and Singing). The study was conducted in a controlled, noise-attenuated environment. All participants used professional-grade studio headphones to ensure they could discern subtle acoustic details and potential artifacts. While certain metrics are common across all tasks, we designed specific dimensions to capture the unique nuances of different audio types. A Likert scale (1â5) was employed for all dimensions. The criteria for each task are summarized in Table 5. A critical challenge in multimodal evaluation is category ambiguity (e.g., a video containing both ambient noise and a specific sound effect). To address this, we implemented a nuanced 1â5 scoring system for Instruction Following rather than a binary âyes/noâ choice. This allows participants to penalize the model less severely when the visual cues are inherently subtle, ensuring a fairer and more stable evaluation of the modelâs intent-alignment capabilities. To minimize cognitive load and ensure consistent judgments across samples, we developed a standardized annotation interface, as illustrated in Figure 17. For each sample, participants are pre- sented with the video together with its generated audio and are asked to evaluate it along four predefined dimensions using a 5- point Likert scale. The interface also provides simple navigation controls, allowing participants to replay the sample and proceed through the evaluation in a structured and efficient manner. 16 V2A VT2A Positive Instruction: "Realistic foley sound synchronized with the video. " Negative Instruction: "music, background music, speech, singing" Positive Instruction: "Realistic foley sound ofa yellow canary calling while flapping its wings inside a metal cage." Negative Instruction: "music, background music, speech, singing" Positive Instruction: "A cappella singing voice synchronized with the video." Negative Instruction: "speech, talking, instrumental only, heavy accompaniment" Positive Instruction: "Singing performance by a female singer with long, wavy brown hair with an expressive, emotionally intense face. Her mouth open mid-note and eyes slightly closed to convey deep feeling, while executing subtle, fluid hand gestures." Negative Instruction: "speech, talking, instrumental only, heavy accompaniment" V2A VT2A V2A VT2A Positive Instruction: "Human speech synchronized with the video." Negative Instruction: "background music, singing, noise, sound effects " Positive Instruction: "Speech by a woman in her 50s with a neutral expression speaking indoors in front of a computer monitor." Negative Instruction: "background music, singing, noise, sound effects " V2A VT2A Positive Instruction: "Musical instrument performance synchronized with the video. " Negative Instruction: "speech, singing, human voice" Positive Instruction: "Instrument performance of a person playing a black accordion with both hands while seated, in a festive setting decorated with orange and yellow balloons and floral arrangements. " Negative Instruction: "speech, singing, human voice" Figure 18: Examples of V2A and VT2A instruction prompts across four audio categories. 17