Paper deep dive
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
Yan-Bo Lin, Jonah Casebeer, Long Mai, Aniruddha Mahapatra, Gedas Bertasius, Nicholas J. Bryan
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 6:20:09 AM
Summary
V2M-Zero is a zero-pair video-to-music generation framework that achieves temporal synchronization by leveraging intra-modal event curves. By computing similarity-based event curves independently for music and video, the model enables training on text-music pairs and inference-time transfer to video without requiring paired cross-modal training data.
Entities (5)
Relation Signals (3)
V2M-Zero → evaluatedon → OES-Pub
confidence 98% · Across OES-Pub, MovieGenBench-Music, and AIST++, V2M-Zero achieves substantial gains
V2M-Zero → usesencoder → DINOv2
confidence 95% · We use MusicFM and DINOv2 by default
V2M-Zero → usesencoder → MusicFM
confidence 95% · We use MusicFM and DINOv2 by default
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-Zero, a zero-pair video-to-music generation approach that outputs time-aligned music for video. Our method is motivated by a key observation: temporal synchronization requires matching when and how much change occurs, not what changes. While musical and visual events differ semantically, they exhibit shared temporal structure that can be captured independently within each modality. We capture this structure through event curves computed from intra-modal similarity using pretrained music and video encoders. By measuring temporal change within each modality independently, these curves provide comparable representations across modalities. This enables a simple training strategy: fine-tune a text-to-music model on music-event curves, then substitute video-event curves at inference without cross-modal training or paired data. Across OES-Pub, MovieGenBench-Music, and AIST++, V2M-Zero achieves substantial gains over paired-data baselines: 5-21% higher audio quality, 13-15% better semantic alignment, 21-52% improved temporal synchronization, and 28% higher beat alignment on dance videos. We find similar results via a large crowd-source subjective listening test. Overall, our results validate that temporal alignment through within-modality features, rather than paired cross-modal supervision, is effective for video-to-music generation. Results are available at this https URL
Tags
Links
- Source: https://arxiv.org/abs/2603.11042v1
- Canonical: https://arxiv.org/abs/2603.11042v1
Trouble viewing inline? Open PDF directly →
Full Text
70,357 characters extracted from source content.
Expand or collapse full text
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation Yan-Bo Lin 1⋆ Jonah Casebeer 2 Long Mai 2 Aniruddha Mahapatra 2 Gedas Bertasius 1 Nicholas J. Bryan 2 1 UNC Chapel Hill 2 Adobe Research Abstract. Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-Zero, a zero-pair video-to-music generation approach that outputs time-aligned music for video. Our method is motivated by a key observation: temporal synchronization re- quires matching when and how much change occurs, not what changes. While musical and visual events differ semantically, they exhibit shared temporal structure that can be captured independently within each modal- ity. We capture this structure through event curves computed from intra- modal similarity using pretrained music and video encoders. By measur- ing temporal change within each modality independently, these curves provide comparable representations across modalities. This enables a simple training strategy: fine-tune a text-to-music model on music-event curves, then substitute video-event curves at inference without cross- modal training or paired data. Across OES-Pub, MovieGenBench- Music, and AIST++, V2M-Zero achieves substantial gains over paired- data baselines: 5–21% higher audio quality, 13–15% better seman- tic alignment, 21–52% improved temporal synchronization, and 28% higher beat alignment on dance videos. We find similar results via a large crowd-source subjective listening test. Overall, our results validate that temporal alignment through within-modality features, rather than paired cross-modal supervision, is effective for video-to-music generation. Results are available at https://genjib.github.io/v2m_zero/ 1 Introduction Generative music is growing in popularity among creators from online influencers on social media platforms (e.g., YouTube, Instagram, TikTok) to professionals in film, gaming, and advertising. Such content creators seek music that both complements their video content and supports fast and flexible control over style and pacing. While recent text-to-music (T2M) methods [1, 15, 21, 24, 40, 57, 63, 73,81,101,110] enable automatic music generation from textual prompts, their outputs are not designed to follow the temporal dynamics of a target video. As ⋆ Work done during an internship at Adobe Research. arXiv:2603.11042v1 [cs.CV] 11 Mar 2026 2Lin et al. Video Features Music Features Training Inference Problem Swap! Music Video DiT Prompt Music Output Prompt (predicted) Synchronized Music Output DiT MusicVideosNo Paired Data!!! Fig. 1: Zero-Pair Video-to-Music Generation Top: Generating music for video commonly requires large-scale collections of high-quality, paired video-music data. Middle: Our V2M-Zero method is trained only on text–music pairs with an ad- ditional music-event curve condition (no video). Bottom: At inference, we swap a music-event curve with aligned video-event curves extracted via off-the-shelf vision models and generate time-synchronized music to match the input video. a result, creators have to manually and delicately edit videos to fit the generated music for synchronization, a tedious and time-consuming process. For instance, in a short product promotion video, musical cues must align with each reveal or motion highlight to create a strong and memorable impression within seconds. This limits the adoption of existing T2M models in real-world content creation. The lack of time-synchronized text-to-music generation motivates video-to- music (V2M) models [13,35,44,47,49,56,87,88,95,97,111], or the task of creat- ing background music that is both temporally and semantically aligned with a given video. Existing V2M methods typically rely on paired video–music datasets curated from online videos, entangling these two aspects of control. Recent work [26,103] explores an alternative direction by using pretrained multimodal large language models (MLLMs) to infer music prompts for video, which are then used as input to a pretrained T2M model. While promising for semantic align- ment, such methods do not explicitly model temporal correspondence between visual and musical events. This motivates a natural question: can we perform time-synchronized video-to-music generation without paired video-music data? To address these limitations, we propose V2M-Zero– a zero-pair video-to- music generation method that generates time-synchronized music from videos without paired video–music data. Our key observation is that synchroniza- tion primarily depends on when change occurs, rather than what changes. V2M-Zero3 Music-event Curve New Objects (Explosion+Logo) Scene Change Human Motion Video-event Curve Fig. 2: Shared Temporal Structure Across Modalities. Real event curves computed from video and music ex- hibit similar temporal patterns across diverse video scenarios. Ground-truth pairs have correlation ≈ 0.6, intro- ducing random offsets degrades this to ≈ 0.2. 1 In practice, video and music synchro- nization often corresponds to (sparse) moments of interest or events over time (e.g., video events of dancing and scene transitions match music events like beats, instrumental and/or dynamic changes). We represent events over time as event curves based on intra-modal similarity, creating structurally comparable tem- poral representations across video and music as shown in Figure 2. Crucially, our design decouples temporal structure from semantic and emotional ground- ing: event curves specify when musical change should occur, while textual condi- tioning determines how it should sound. Through lightweight fine-tuning of a pre- trained T2M model with added music event curves, we enable test-time transfer by simply substituting video event curves at inference. For text conditioning, we ex- tract visual captions from video and summarize them with an LLM to capture the mood [26], achieving robust adaptation to diverse and complex scenes. 2 Our contributions are threefold. First, we propose V2M-Zero, the first zero- pair framework for time-synchronized video-to-music generation. Our key insight is that event curves derived from intra-modal similarity yield structurally com- parable temporal representations, enabling transfer from music conditioning dur- ing training to video conditioning at inference with only lightweight fine-tuning (192-768 GPU hours) of a pretrained model. Second, through extensive abla- tions, we validate the flexibility of our design: V2M-Zero adapts to different video domains simply by selecting appropriate visual encoders (e.g., foundation models for general video, motion trackers for dance) without retraining, and we analyze design choices for mitigating the modality gap between music and video representations. Lastly, we demonstrate state-of-the-art results across three di- verse benchmarks. Our zero-pair approach outperforms paired-data supervised methods by substantial margins on objective metrics: 5–21% better in music quality, 13–15% in semantic alignment, 21–52% in temporal synchronization, and 28% in beat alignment on dance videos. Furthermore, we find similar results via a large crowd-source subjective listening test. 1 Correlation is computed in a 1-second window around each scene cut in a video. Correlation degrades when the music is time-shifted relative to the video. 2 Demos with fixed event-curves but varying text inputs are in the demo page. 4Lin et al. 2 Related Work 2.1 Text-to-Music Generation General text-to-audio generation has advanced significantly in the past few years [4, 19, 20, 30, 36, 50–52, 62, 72, 83, 84, 91, 99]. Text-to-music generation has followed a similar trend [1,15, 21,24,40,57, 63,73,81,101, 110]. Both music and audio generation methods typically fall into one of two modeling paradigms. The first, autoregressive (AR) modeling [1,15,36,40,110], operates on discrete audio tokens produced by a neural audio codec [16,37,102]. A causal transformer pre- dicts each token conditioned on the previously generated tokens and the input text prompt, and the resulting token sequence is decoded back to a waveform by the codec. The second, latent diffusion models (LDMs) [21,23,50,51,61,63], learns a denoising process over continuous latents conditioned on text. These employ a similar kind of neural audio codec with continuous latents [20]. While recent text-to-music models [1,15,21,24,40,57,73,81,110] can effectively capture high-level semantics, such as genre, mood, and instrumentation, it remains chal- lenging to align music with fine-grained visual events. This limitation motivates the study of video-to-music generation with tighter temporal correspondence and control between video and music. 2.2 Video-to-Music Generation Building on video-to-audio research [8–12,18,38,55,58,60,66,79,93,104], recent video-to-music methods aim to generate soundtracks aligned with visual rhythm and motions (i.e., dance videos) [41, 45, 77, 80, 100, 106–108]. Early approaches relied on symbolic data such as MIDI or ABC notation [17,32,46,86,109], but these datasets were limited in scale and expressivity. More recent works train on paired internet videos accompanied by music tracks [13, 35, 44, 47, 49, 53, 56, 70,78,87–89,95,105,111]. However, such internet data is often noisy, containing vocals, imperfect mixing, or potential copyright issues. These limitations hinder the development of high-fidelity models and motivate the exploration of unpaired or zero-shot paradigms that leverage independent music and video sources. 2.3 Video-to-Music Generation via Prompting Without paired video-music training data, recent methods [31,42,75,92,97,103] bridge the video and music domains through prompts. Specifically, they translate visual content into text prompts (via LLM reasoning or learned prediction) and then generate music using off-the-shelf text-to-music models. While this strategy effectively captures high-level semantics (e.g., genre, mood), it struggles with fine-grained temporal alignment because such prompts lack the expressiveness to specify timing and dynamics. In contrast, V2M-Zero directly conditions on music or video event curves, enabling temporal synchronization with scene cuts, motion patterns, and dance rhythm. V2M-Zero5 Visual Encoder Music Encoder Training Inference Music Video Music Output Prompt: ... Cosine Similarity Inv. Smooth ... Cosine Similarity Inv. Smooth +Noise “4/4 stringostinato in D minor” Music Captioner Predict An epic cinematic orchestral ...” Noise DiT DiT Time-Synchronized Music Output Swap ! Video-event curve Music-event curve Speech Music Prompt Fig. 3: Method Overview Top: During training, V2M-Zero learns a rectified-flow diffusion process conditioned on text prompts and a music-event curve derived from intra-music similarity. Bottom: At inference, music conditioning is swapped with a video-event curve based on framewise similarity, enabling zero-pair, time-synchronized video-to-music generation. For semantic alignment, a text prompt is predicted from the video and speech using the music captioner, Vibe [26], without any joint-training. 3 Technical Approach We fine-tune a pretrained T2M latent rectified flow model by augmenting it with temporal event curve conditioning, enabling test-time transfer to video at inference. Our key insight is that temporal change, measured via intra-modal similarity, captures domain-agnostic event structure. Events such as beat on- sets in music or scene cuts in video both manifest as local dissimilarity between consecutive temporal segments. By standardizing these dissimilarity signals, we obtain structurally comparable representations across modalities. During train- ing, the model is conditioned on both text prompts and music-event curves. At inference, we replace the music-event curve with a video-event curve for zero- paired, time-synchronized generation without any model retraining – alleviating the need for paired video-music data while generating temporally aligned out- puts. An overview is shown in Figure 3. 3.1 Preliminaries Music Autoencoder. We use a pretrained music autoencoder [7] to encode stereo music waveforms into continuous latents x 0 ∈ R d×l , where d is the latent dimension and l is the temporal length. These latents represent the low-level 6Lin et al. generative representation used by our rectified flow model. Separately, we extract high-level semantic features from pretrained encoders to compute event curves. Rectified Flows. Given a text condition c and a music latent x 0 sampled from the empirical data distribution P D , a rectified flow model learns a velocity field f θ (· ) that transports samples from a Gaussian prior N (0,I) to P D . The timestep variable t∈ [0, 1] interpolates between noise (t = 1) and data (t = 0) via the linear path x t = tε + (1− t)x 0 , whereε∼N (0,I). The model f θ (· ) is trained to predict the velocityε− x 0 via: min θ E x 0 ,ε,t,c ∥ (ε− x 0 )− f θ (x t ,c,t)∥ 2 2 .(1) At inference, samples are generated by solving the ODE dx t = −f θ (x t ,c,t)dt from t = 1 to t = 0 with 96 sampling steps and classifier-free guidance [28]. Model Architecture. We use a pretrained Diffusion Transformer (DiT) archi- tecture [67] f θ (· ) with cross-attention text conditioning c following standard T2M models [20,21,23,50,51,61]. 3.2 Temporal Event Curves We construct event curves or 1D temporal signals that capture when and how much change occurs (sometimes denoted as novelty curves). Our procedure is identical for music and video, enabling direct substitution at inference. Perspectives on Event Curves. Our event curve formulation connects to es- tablished uses of self-similarity across modalities. In music, self-similarity has been widely used to analyze musical structure and rhythm [22, 68] and is also considered a generalization of the concept of autocorrelation. In video, similar approaches segment footage and detect natural boundaries such as shot tran- sitions [39, 76] Mathematically, our approach is equivalent to extracting one or more off-diagonal bands of the self-similarity matrix. In that sense, our formula- tion leverages the natural structure of self-similarity to enable zero-shot transfer. Features for Event Curves. To compute event curves, we begin by extracting temporal feature sequences using pretrained encoders. For music (training only), we apply a music encoder to obtain f m ∈ R d m ×l m , where l m is the temporal length and d m is the feature dimension. For video (inference only), we encode each frame with a visual encoder and spatially pool to obtain f v ∈ R d v ×l v , where l v is the number of frames and d v is the feature dimension. These high-level semantic features are used solely to compute event curves. Computing Event Curves. Given a feature sequence f ∈ R d f ×l f (either f m or f v ), where f k denotes the k− th temporal feature-vector, we measure temporal change via the cosine similarity between consecutive vectors: s k = f k · f k+1 ∥f k ∥f k+1 ∥ , k = 1,...,l f −1.(2) We obtain a dissimilarity sequence via a k = 1− s k and set A = a k . Higher values indicate stronger temporal change, capturing events such as beat onsets in music or scene cuts in video. V2M-Zero7 Mitigating the Modality Gap. To enable zero-shot transfer from music to video, we apply three operations that align event curves across modalities. First, we standardize any A to have zero mean and unit variance: ̄a k = a k − μ(A) σ(A) .(3) Without standardization, music and video event curves have different scales and offsets, creating a distribution shift when swapping at inference. Second, we resample to length l (matching the temporal dimension of x 0 ). Third, we apply temporal smoothing with a Hann window to suppress modality-specific details while preserving the larger structure, yielding the final event curve: e = Smooth Resample ̄ A,l ∈ R l .(4) During training, we compute e m from music features. At inference, we compute e v from video features. Our design is agnostic to the choice of feature encoders. We use MusicFM [94] and DINOv2 [65] by default but explore alternative en- coders in Section 5.3. 3.3 Rectified Flow Model Fine-Tuning We incorporate event curves into a pretrained text-conditioned rectified flow model via fine-tuning. Event Curve Conditioning via Concatenation. We inject the event curve e by concatenating it as an additional channel to the rectified flow latent: e x t = x t , e ∈ R (d+1)×l , (5) where [·,·] denotes channel-wise concatenation. This approach is simple and parameter-efficient, only requiring additional parameters in the input projection layer of the DiT. For simplicity, we only use a single event curve, though the conditioning can naturally incorporate multiple curves from different temporal scales. Fine-Tuning Objective. We initialize from a pretrained T2M model, add the conditioning signal, and fine-tune it with music-event curve conditioning via: min θ E x 0 ,ε,t,e m ,c (ε− x 0 )− f θ e x t ,c,t 2 2 , (6) where we explicitly denote e m to emphasize that during training, event curves are computed from music. Given a music clip during training, we extract music encoder features f m and compute the music event curve e m via the procedure in Section 3.2. The text condition c is obtained from ground-truth music descrip- tions in the pre-training data. Fine-tuning allows our model to learn to generate music conditioned on both text and temporal events. 8Lin et al. 3.4 Zero-Pair Inference via Event Curve Swapping At inference, we perform test-time transfer to video-to-music generation by swap- ping our event curve conditioning signals while keeping all model weights, θ, fixed. Given an input video, we extract visual encoder features f v and compute the video event curve e v using the identical procedure applied to music during training. For text conditioning, we use an LLM [26] to generate and summarize video captions into a music-appropriate text prompt c. We then generate mu- sic using standard flow-model inference and add the resulting generated audio track to the input video. Note, our music and video event curves are designed to mitigate the domain gap, so this swapping does not require retraining. 3.5 Implementation Details Architecture. We adopt an internal, pretrained T2M model similar to [3, 64] for fine-tuning. The audio autoencoder is similar to [3, 5, 6, 20] and compresses stereo 44.1 kHz waveforms into continuous latents (d=64) at 12.3 Hz, yielding 394 frames for 32-second clips. The rectified flow model uses the DiT architecture with≈ 1 billion parameters: 16 layers, hidden dimension 2048, and feed-forward dimension 8192. Text conditioning is implemented via cross-attention layers us- ing Gemma-3B [85] embeddings. Feature Encoder. We use MusicFM [94] and DINOv2-L [65] to extract music features f m and video features f v by default. We explore a comprehensive ablation of encoder choices, including an audio autoencoder, AVSiam [48], V-JEPA [2], and CoTracker [33] in Section 5.3. Training. We initialize from a pretrained text-to-music model, adding 2048 parameters to input and project our event curve conditioning, and fine-tune on ≈ 25k hours of licensed instrumental music-text pairs. Fine-tuning is lightweight, requiring only 2–4 days on 4–8 A100 GPUs using 32-second clips (192-768 GPU hours). We use the AdamW optimizer with a learning rate 10 −4 and apply classifier-free guidance with 10% condition dropout. Cumulatively, this fine- tuning scheme enables video-to-music generation without paired video-music data or training from scratch given a pretrained model. 4 Experimental Setup We conduct experiments on three diverse benchmarks to validate our zero-pair approach. This section describes our datasets, metrics, and baselines. 4.1 Evaluation Data We test on datasets spanning general, cinematic, and dance: – OES-Pub [35] is the public evaluation split of the Open Screen Sound- track Library (OSSL), containing 115 public-domain movie clips paired with royalty-free music. Each clip is ≈ 30 seconds and includes human-annotated music prompts. V2M-Zero9 – MovieGenBench-Music [69] is the music subset of MovieGenBench, con- sisting of 527 video-music pairs with sound effects across diverse video types. Each clip is ≈ 10 seconds and includes a music prompt. – AIST++ [43, 90] is a street dance dataset with copyright-cleared dance music, consisting of 20 video–music pairs across 10 dance genres. Each clip is ≈ 7 seconds and includes BPM. 4.2 Evaluation Metrics Following [49, 87, 88, 111], we measure: audio fidelity, semantic alignment, and temporal synchronization. For each dataset under test, we customize the metrics used to highlight the critical aspects of the specific video-to-music generation task. OES-Pub & MovieGenBench-Music: – Fréchet Audio Distance (FAD) [34]: audio fidelity via distributional dis- tance between reference and generated music in VGGish [27] space. We de- note the use of a separate reference set [59] via *. Lower is better. – CLAP Score [96]: semantic alignment via cosine similarity between gener- ated music and text prompts in CLAP space. Higher is better. – Scene Cut Hit (SCH) 3 : aperiodic temporal alignment. A hit occurs when a beat falls within ±100 ms of a scene cut, computed as hits/total cuts. Higher is better. – Human Evaluation: measures subjective preference for music quality and video synchronization. Given two music tracks for the same video, raters answer: 1) Which has better quality? and 2) Which better synchronizes with visuals? We collected 1403 valid votes from the crowd-source platform Ap- pen, evenly sampled across datasets and randomly sampled across baselines We report win-rates with pairwise t-tests vs. [54,87,88,95,103,111]. Higher is better. AIST++: – Beat Coverage (BCS), Beat Hit Score (BHS), F1: periodic rhythm alignment [45]. BCS and BHS are recall and precision of music beats relative to motion beats, and F1 is their harmonic mean. Higher is better. – Temporal Deviation (TD): tempo difference from ground truth. We use .2s tolerance (reduced from 1.0s) for perceptual meaningfullness [45]. Lower is better. 5 Results and Analysis 5.1 Comparison with the State-of-the-Art V2M-Zero demonstrates strong performance across all benchmarks, outper- forming both paired/unpaired baselines despite using no paired video-music 3 Pseudocode in appendix. 10Lin et al. Table 1: Comparison with State-of-the-Art. We evaluate all methods [54,87,88, 95,103,111] on OES-Pub [35] and MovieGenBench-Music [69] in audio quality (FAD), high-level semantic alignment (CLAP), and our proposed temporal metric (SCH). V2M-Zero outperforms all competing methods across all datasets and metrics. Note, OES-Pub reference data includes non-musical sounds, so we use SongDescriber [59] for the reference distribution for FAD results. OES-Pub.MovieGenBench-Music Audio Quality High-Level Alignment Low-Level Matching Audio Quality High-Level Alignment Low-Level Matching MethodV–M PairsFAD*↓ CLAP↑ SCH↑FAD↓ CLAP↑ SCH↑ M 2 UGen [54]✔ 36.7h6.670.070.355.840.020.24 GVMGen [111]✔ 147h6.250.180.353.960.060.48 MTCV2M [95]✔ 147h5.440.200.374.020.160.29 VidMuse [88]✔ 18000h10.40.160.402.980.040.47 AudioX [87] ✔ 15793h7.460.190.332.820.080.30 SONIQUE [103]✘0h6.800.090.276.470.160.21 V2M-Zero (Ours)✘0h4.950.230.612.680.180.58 training data. We analyze results across audio fidelity, semantic alignment, and temporal synchronization, highlighting key insights on cross-domain generaliza- tion and the utility of event-curves. General & Cinematic Video to Music Generation. Table 1 shows results on OES-Pub and MovieGenBench-Music using DINOv2 [65] as the vi- sual encoder. V2M-Zero achieves the best audio quality on both benchmarks (i.e., FAD*: 4.95 on OES-Pub; FAD: 2.68 on MovieGenBench). Notably, our method exhibits the most consistent performance across datasets, while paired baselines show substantial variation (e.g., VidMuse: 10.4 → 2.98). Interestingly, VidMuse, trained on 18,000 hours of paired data, achieves the worst FAD on OES-Pub but the best among paired methods on MovieGenBench, suggesting dataset-specific overfitting rather than robust generalization. We compute FAD* on OES-Pub using SongDescriber [59] as the reference distribution, since the original OES-Pub audio contains speech, sound effects, and background noise. MovieGenBench uses generated music references, which generally yield lower absolute FAD scores than real recordings. For semantic alignment (CLAP), paired-data baselines perform reasonably on OES-Pub but exhibit noticeable degradation on MovieGenBench-Music. For instance, AudioX achieves 0.08 on MovieGenBench-Music, compared to 0.19 on OES-Pub. This suggests limited generalization to MovieGenBench’s generated videos, which differ from the real paired data used during training. Interestingly, SONIQUE is the only baseline that improves on MovieGenBench (i.e., 0.09 → 0.16), likely because its pure LLM-based prompting approach is less sensitive to domain shift in the visual content. In contrast, V2M-Zero demonstrates robust performance across both datasets (i.e., 0.23 on OES-Pub, 0.18 on MovieGenBench-Music), indicating strong cross-domain generalization while maintaining substantially higher ab- solute CLAP scores than all baselines. V2M-Zero11 Table 2: Human Evaluation. Pairwise win rates of our model against each baseline (1403 ratings, Bonferroni-corrected t-tests, p < 0.0083). † above chance but not signif- icant after correction. Scene cut subset comprises 67% of all votes. (a) All Videos Win Rate of Ours ↑ BaselineMusic Quality Temp. Align. AudioX72.36%60.36% GVMGen67.00%67.00% M 2 UGen67.45%66.51% MTCV2M 59.72%59.72% SONIQUE77.16%73.60% VidMuse68.85%53.77% † Average68.76%63.49% (b) Scene Cut Videos Win Rate of Ours ↑ BaselineMusic Quality Temp. Align. AudioX73.17%61.46% GVMGen70.29%67.39% M 2 UGen 67.88%66.42% MTCV2M65.35%66.14% SONIQUE81.54%78.46% VidMuse 74.63%59.51% Average71.14%66.56% For temporal alignment (SCH), methods trained on paired video-music data achieve strong performance, ranging from 0.33–0.40 on OES-Pub and 0.24– 0.48 on MovieGenBench-Music. In comparison, SONIQUE, which uses a pure LLM-based prompting approach, achieves lower temporal alignment (i.e., 0.27 on OES-Pub, 0.21 on MovieGenBench-Music). This demonstrates the bene- fit of learning from paired data for capturing temporal structure. However, V2M-Zero achieves strong temporal alignment (i.e., 0.61 on OES-Pub, 0.58 on MovieGenBench-Music), showing that explicit event-curve conditioning can surpass paired supervision for fine-grained synchronization. Human Evaluation. We analyze our 1403 pairwise human ratings across six baselines and two questions via multiple pairwise t-tests with Bonferroni correction [29]. All win-rates reported in Table 2 show our win-rate (higher is better). We show results across all videos in Table 2(a) and videos with a scene cut in Table 2(b). Overall, our method is preferred with a > 50% win-rate. However, statistical significance is a bit varied, and we find it is tightly coupled with whether the rated videos contain a clear event, like a scene cut. On general videos, our approach has a statistically significant win-rate over all baselines in music quality. We see a similar trend in temporal alignment, where our method has a statistically significant win-rate against 5/6 baselines. To understand when and why V2M-Zero is preferred, we test scenes with a scene-cut in Table 2(b), a clear event for raters to anchor their evaluation on. We observe that win-rates either maintain or score more significant. We find that V2M-Zero is preferred over paired-data baselines and that this preference increases with an anchoring visual event. 5.2 Generalization Across Video Types and Models Dance Video to Music Generation. To evaluate our approach on content with precise, tightly coupled temporal dynamics, we test on AIST++ [43, 90], 12Lin et al. Table 3: Generalization Across Video Types and Model Implementations. (a) Eval. on [43, 90] using beat consistency (BCS, BHS, F1) and temporal deviation (TD). (b) V2M-Zero on a public text-to-music model evaluated on OES-Pub. (a) Dance Video Results on AIST++. MethodBCS↑ BHS↑ F1↑ TD↓ CMT [17]0.3368 0.1515 0.2090 21.74 CDCD [107] 0.4233 0.2151 0.2852 19.25 LORIS [100]0.3721 0.3371 0.3537 17.80 MDM [82] 0.3798 0.4185 0.3982 22.96 Text Inv. [45]0.4761 0.4398 0.4572 20.34 V2M-Zero (Ours)0.5818 0.6274 0.5856 12.24 (b) Cross-Architecture General- ization (OES-Pub). MethodFAD*↓ CLAP↑ SCH↑ SD-Audio-ctrl [14] 4.86 0.18 0.28 +V2M-Zero(Ours) 4.13 0.17 0.38 a dance dataset where video and music are more densely synchronized. Ta- ble 3(a) shows results comparing against methods explicitly designed or trained for dance-to-music generation. Rather than requiring dance-specific paired train- ing, our method adapts by using a domain-appropriate visual encoder. Specifi- cally, we use CoTracker [33], a point-tracking model that captures fine-grained motion trajectories suited to dance. With this encoder, V2M-Zero achieves strong performance across all rhythm metrics (0.5818/0.6274/0.5856/12.24), out- performing specialized methods. Notably, improvements are most pronounced for BHS and TD, suggesting that event curve conditioning achieves precise and ac- curate motion-rhythm correspondence. This highlights a key advantage of our approach: by decoupling event-curve extraction from music generation, the same model can flexibly adapt to diverse video domains without retraining. Generalization to Public Text-to-Music Models. To confirm the gener- ality of V2M-Zero, we evaluate with a pre-trained publicly available text-to- music model, Stable-Audio-ControlNet [14]. This model is trained with audio RMS curves, which we swap for video-event curves at inference time. Table 3(b) shows comparable audio quality (FAD*) and semantic alignment (CLAP), but improved Scene Cut Hit (SCH) from 0.28 to 0.38. The improvement in temporal alignment, via zero-shot integration, indicates that event-curve conditioning is model-agnostic and not tied to a specific backbone. 5.3 Ablation Studies 5915233139495563 Kernel Size 3 4 5 6 7 8 FAD* FAD* Scene Cut Hit Acc. 0.2 0.3 0.4 0.5 0.6 0.7 Scene Cut Hit Acc. Fig. 4: Impact of Smoothing Kernel Size. Larger kernerls im- prove audio quality (FAD*) but temporal alignment (SCH) has an optimal point on OES-Pub. We systematically ablate four design axes: (i) kernel size for modality gap mitigation, (i) encoders for event-curve extraction, (i) domain-specific visual encoders, and (iv) LLM selection for prompt generation. Mitigating Modality Gap. Music-event curves (training) and video-event curves (in- ference) differ in temporal granularity, creat- V2M-Zero13 Table 4: Encoder Selection for Event Curves. Impact of different music encoders (training) and visual (inference) encoders on OES-Pub using FAD, CLAP, and SCH. MusicFM + DINOv2 achieves the best balance across metrics. Music Enc. (Training) Visual Enc. (Inference) FAD*↓ CLAP↑ SCH↑ AVSiam [48]4.52 0.19 0.35 VAE [7]V-JEPA 2 [2] 5.130.18 0.41 VAE [7]DINOv2 [65] 4.770.16 0.31 MusicFM [94] V-JEPA 2 [2] 5.020.18 0.48 MusicFM [94] DINOv2 [65] 4.95 0.23 0.61 ing a modality gap that can degrade zero-shot transfer. We mitigate this via Hann-window smoothing applied to both modalities. In Figure 4 we find that in- creasing the kernel size from 9 (.7 seconds) to 63 (5 seconds) improves FAD from 8.17 to 3.12. However, excessive smoothing blurs fine-grained events, causing SCH to degrade from 0.61 to 0.27. This highlights a trade-off: stronger smooth- ing improves audio quality but weakens temporal alignment. We pick a kernel size of 31 as it balances these two competing objectives. Encoder Architecture for Event Curves. Table 4 compares three en- coder design paradigms: shared, reconstruction, and self-supervised encoders. Shared Encoders. We test whether shared-weight audio-visual encoders, like AVSiam [48], can reduce the modality gap by embedding both modalities in a common space. AVSiam achieves the best FAD (4.52) among all configura- tions, suggesting improved distribution matching. However, temporal alignment degrades significantly (SCH: 0.61→ 0.35), likely due to AVSiam’s lower special- ized capacities in both the audio and visual domains compared to foundation models like DINOv2 [65] and a fundamental difference on how music and au- dio are aesthetically matched to music. We hypothesize that large-scale shared encoders could close this gap as a promising future direction. Music Encoder. The music encoder has the strongest impact on performance. We initially use the same encoder as the latent diffusion VAE, for simplicity. Us- ing a self-supervised model like MusicFM [94] improves temporal synchronization (SCH: 0.31 → 0.61, with DINOv2), while also improving audio quality (FAD: 4.95 → 4.77). MusicFM’s more semantic representations better correlate with visual event patterns, enabling more reliable curve alignment during inference. Video Encoder. We compare self-supervised visual encoders: V-JEPA [2] and DINOv2 [65]. DINOv2 yields slightly better audio quality across music encoders (FAD: 4.77 vs. 5.13 with VAE; 4.95 vs. 5.02) with MusicFM [94]). However, align- ment results are mixed: V-JEPA outperforms with VAE (SCH: 0.41 vs. 0.31), while DINOv2 pairs best with MusicFM (SCH: 0.61 vs. 0.48). Overall, video encoder choice has minimal impact; alignment quality depends more critically on the music encoder. We use MusicFM + DINOv2 as our default configuration. 14Lin et al. Table 5: LLM Selection for Music Prompts. Impact of different LLMs for gen- erating music prompts from video on OES-Pub. V2M-Zero uses Vibe [26], based on Gemma-4B [85]. Music CaptionerFAD*↓ CLAP↑ SCH↑ Qwen3-4B [98]4.98 0.23 0.58 Llama-3.2-3B [25]5.020.21 0.60 Gemma-4B (default) [85] 4.95 0.23 0.61 Domain-Specific Visual Encoders. Our event-curve framework enables performance gains via inference-time encoder selection. We validate this on dance videos (AIST++), where motion is tightly synchronized to a single object. Our default configuration with DINOv2 outperforms all the competing methods (BCS: 0.5522, BHS: 0.5748, F1: 0.5750, TD: 17.23). We hypothesize that Co- Tracker [33], a point-level motion tracker, is more suitable for this task. This is true and improves all metrics (BCS: 0.5818, BHS: 0.6274, F1: 0.5856, TD: 12.24) in Table 3. This shows that V2M-Zero can be customized at inference by picking a domain-specific encoder. LLM Selection for Music Prompt Generation. Table 5 compares three LLMs for generating music prompts from video: Gemma-4B [85], Qwen3-4B [98], and Llama-3.2-3B [25]. Event curves are fixed, and we ablate the prompts. Dif- ferences across LLMs are minimal: FAD, CLAP, and SCH vary by less than 5%. Gemma-4B achieves marginally better results (FAD: 4.95, CLAP: 0.23, SCH: 0.61). We conclude that LLM choice has a negligible impact on performance. Any modern LLM suffices for semantic guidance, assuming a similar task setup. 6 Conclusions We introduce V2M-Zero, a zero-pair framework for time-aligned video-to-music generation that bridges modalities through temporal structure rather than paired supervision. Our observation is that while musical and visual events differ se- mantically, they exhibit similar temporal patterns when embedded by pretrained music and video encoders. We exploit this property using event curves, signals computed from consecutive frames or segment dissimilarity within each modal- ity’s feature space. These curves act as a flexible conditioning signal: a text-to- music model fine-tuned on music-event curves directly accepts video-event curves at inference time, requiring no architectural changes or cross-modal data. Our ablations reveal actionable insights on modality gap mitigation, encoder selec- tion, and domain-specific encoder specialization. On OES-Pub, MovieGenBench- Music, and AIST++, V2M-Zero consistently achieves state-of-the-art results, outperforming paired-data methods in audio quality, semantic alignment, and temporal synchronization. Overall, our results show that temporal alignment through intra-modal features is a viable alternative to paired data. V2M-Zero15 For future work, we plan to perform a deeper qualitative investigation of real, high-quality matched video-to-music data to better understand the artistic stylizations of event synchronization, investigate video-to-music generation in the context of low-resource video-to-music data pairs as opposed to fully zero-pair fine-tuning, and further improve the cross-domain curve alignment to further mitigate the modality gap. Acknowledgments This work was supported by Laboratory for Analytic Sciences via NC State University, ONR Award N00014-23-1-2356, NIH Award R01HD11107402, and Sony Focused Research award. A Appendix Overview Our appendix consists of: 1. Additional Implementation Details. 2. Additional Quantitative Results. 3. Qualitative Results on Event Curves. Algorithm 1 Implementation Details of Scene Cut Hit. 1 import torchaudio 2 import librosa 3 from scenedetect import detect, AdaptiveDetector 4 5 def process_video(video_path, tolerance=0.1): 6 # Detect scenes 7 scene_list = detect(video_path, AdaptiveDetector()) 8 scene_cut_times = [] 9 for i, (start, end) in enumerate(scene_list): 10 if i < len(scene_list) - 1: 11 scene_cut_times.append(end.get_seconds()) 12 13 if len(scene_cut_times) == 0: 14 return None 15 16 # Load audio 17 x, sr = torchaudio.load(video_path) 18 x = x.mean(0).numpy() 19 20 # Beat/onset extraction 21 tempo, beats = librosa.beat.beat_track(y=x, sr=sr) 22 onsets = librosa.frames_to_time(beats, sr=sr) 23 24 # Hit-rate computation 25 hits = 0 26 for c in scene_cut_times: 27 for o in onsets: 28 if abs(o - c) <= tolerance: 29 hits += 1 30 break 31 32 return hits / len(scene_cut_times) 16Lin et al. Algorithm 2 Music Prompt Generation from Video [26] Require: Video V , Automatic Speech Recognition A, Vision-Language Model L v , Large Language Model L 1: Extract transcript: T ←A(V ) 2: Initialize caption set: C ← ∅ 3: for each sampled frame f i from V do 4: Obtain visual description: C i ←L v (f i ) 5:C ← C∪C i 6: end for 7: Generate visual summary: S ←L(C) 8: Generate music prompt: P ←L(T, S) 9: return P B Additional Implementation Details Scene Cut Hit (SCH). To evaluate the temporal consistency between gen- erated music and the underlying video dynamics, we design and use the SCH metric. The intuition is as follows: When a scene change occurs in the video, well-aligned background music should exhibit a corresponding onset or notice- able rhythmic change. This principle reflects common filmmaking conventions, where directors often synchronize cuts with musical beats to enhance pacing and emotional impact. The SCH metric is conceptually similar to beat matching metics [45] but describes aperioidic visual events. Given a video, we first detect its scene cuts using the PySceneDetect, obtaining a set of scene-cut timestamps that serve as temporal anchors. We then extract the audio track and compute musical onsets using the beat–tracking algorithm provided in Librosa. While more sophisticated onset detectors are available, beat tracking offers a repro- ducible, robust and musically meaningful proxy for rhythmic emphasis and per- cussion peaks. For each detected scene cut, we count a hit if at least one musical onset falls within a small temporal tolerance (we use ±0.1 seconds). The final SCH score is the ratio of matched cuts to all scene cuts. The complete procedure is summarized in Algorithm 1, which outlines the pipeline from scene detection to music onset extraction and the computation of the overall hit rate. VLLM for Music Prompts. To obtain high-level semantic music prompts that reflect the mood/energy/etc. of a video, we adopt the Vibe framework [26] with modern multimodal components. Given an input video, we first obtain its transcript using the Whisper [71] ASR model, which provides a robust textual representation of the spoken content. Dialogue and narration often contain emo- tionally or contextually important cues, making the transcript an essential input to the music–prompt generation process. In parallel, we sample video frames and extract frame-level visual descriptions using Gemma-4B [85]. These descriptions capture high-level semantic attributes such as environment, actions, character states, and affective cues. Since indi- vidual captions may be redundant, we aggregate them using Gemma-4B, which summarizes the set into a compact, coherent representation of the entire video. V2M-Zero17 Table 6: Comparison with a Text-Only Baseline. We compare V2M-Zero against the text-to-music model used for ours finetuning, without event-curve con- ditioning, on OES-Pub. MethodFAD↓ CLAP↑ SCH↑ Text-only 3.63 0.23 0.35 V2M-Zero 4.950.23 0.61 Table 7: Comparison with a Large-Scale Open-Source Model. We compare V2M-Zero with HunyuanVideo-Foley [74] on OES-Pub. MethodFAD*↓ CLAP↑ SCH↑ HunyuanVideo-Foley [74] 15.02 0.124 0.36 V2M-Zero4.95 0.23 0.61 Finally, the LLM conditions jointly on the transcript and the visual summary to produce the final music prompt, which describes mood, instrumentation, in- tensity, and emotional character. This process encourages prompts that capture multimodal semantics and better align with creators’ intuitive descriptions. The full procedure is provided in Algorithm 2. C Additional Quantitative Results. Comparison with the text-only baseline. In Table 6, we compare V2M- Zero with a text-only baseline, i.e., the same text-to-music backbone, check- point, and predicted music prompts without event-curve conditioning, on OES- Pub. The results show that event-curve conditioning substantially improves tem- poral alignment, increasing SCH from 0.35 to 0.61, while preserving semantic consistency, as both models achieve the same CLAP score of 0.23. Although the text-only baseline achieves similar music quality, its much weaker SCH suggests that text guidance alone is insufficient for accurate synchronization. Overall, these results suggest that event curves improve synchronization by injecting ex- plicit temporal structure, while leaving semantic alignment largely unchanged. Comparison with a Large-Scale Open-Source Model. In Table 7, we com- pare V2M-Zero with HunyuanVideo-Foley, a large-scale open-source model de- signed for general audio generation rather than music generation, which has not been evaluated on video-to-music benchmarks. Since HunyuanVideo-Foley generates music but also speech and environmental sounds, it performs substan- tially worse across all three metrics: FAD* (15.02 vs. 4.95), CLAP (0.124 vs. 0.23), and SCH (0.36 vs. 0.61). However, this comparison should be interpreted with caution, as HunyuanVideo-Foley is not designed for music generation and is therefore not well-suited to this setting or our music-focused evaluation protocol. Temporal Alignment Analysis. Since the core idea of V2M-Zero is based on event curves, we further investigate how event-curve distances relate to low- 18Lin et al. Table 8: Event-Curve Fréchet Distance Comparison. M evaluates generated vs. ground-truth music event curves. M+V compares concatenated music–video event curves. M-V compares the distribution of generated music and ground-truth video. M|V evaluates conditional music event distributions given video. We additionally report human preference for temporal alignment. OES-PubMovieGenBench-MusicHuman Eval MethodM M+V M-V M| VM M+V M-V M| VTemporal Align. M 2 UGen [54]6.29 6.96 18.26 1.331.96 1.99 24.76 1.52Ours wins GVMGen [111]3.45 4.04 10.10 0.481.31 1.34 20.09 1.02Ours wins MTCV2M [95]4.33 4.53 10.90 0.921.53 1.75 21.03 0.98Ours wins VidMuse [88]3.25 3.86 12.40 0.794.27 4.29 40.89 4.02Not significant AudioX [87]4.13 4.63 1.75 1.931.52 1.54 19.70 1.18Ours wins SONIQUE [103]3.13 3.75 8.50 0.1511.10 11.12 6.33 9.75Ours wins V2M-Zero (Ours)5.34 5.87 1.20 3.093.27 3.29 27.90 1.75N/A level temporal alignment quality. Specifically, since our method is designed to match these event curves, we attempt to quantify and understand the distribu- tional behavior of this tight synchronization. In Table 8, we report four variants of Fréchet Distance computed between different event-curve distributions: (1) M, generated to ground-truth music event curves, (2) M+V, concatenated gen- erated music-video event curves and ground-truth music-video event curves. (3) M-V, generated music event curves to video event curves, and (4) M|V, gener- ated music event curves and ground-truth music event curves, both conditioned on video event curves. M|V uses a block partitioned Gaussian to represent the music given video conditional distribution. The music event curves are extracted by musicfm [94], and the video event curves are extreacted by DINOv2 [65] We compare these additional quantitative results with human judgments of temporal alignment following the protocol used in the main paper. Interestingly, the human preference results (rightmost column of Table 8) do not correlate strongly with any of the event-curve distances. Models that achieve low Fréchet distances are not necessarily preferred by human raters, and vice versa. This mismatch suggests an important insight: event curves may be highly suitable as a generative intermediate representation, but they may not be ap- propriate as an evaluation metric. Event curves are good at capturing dense, continuous temporal dynamics that help guide generation, yet humans appear to assess alignment based on sparse, salient temporal moments (e.g., impactful transitions or rhythm–scene coincidences), not global curve similarity. As a re- sult, metrics that rely purely on curve distributions reward continuous structural alignment, while human perception is more selective and non-uniform over time. A second contributing factor is that current evaluation datasets (OES-Pub and MovieGenBench-Music) are not specifically curated for temporal synchro- nization assessment. Many clips lack strong or consistent rhythmic structure, limiting the signal available to curve-based metrics. V2M-Zero19 050100150200250 Temporal Position 1.5 1.0 0.5 0.0 0.5 1.0 1.5 2.0 Normalized Dissimilarity Video-event and Generated Music-event Curves (Truncated) Video 1 Generated Music 1 Video 2 Fig. 5: Example event curves with different temporal dynamics. The blue solid curve corresponds to a video with frequent scene cuts, while the orange curve corresponds to a video with slower visual motion, showing distinct temporal structures. This supports our design choice of using event curves to represent relative timing, while text provides complementary semantic guidance. In sum, these observations indicate that while event curves provide a powerful foundation for generation, but the distributional distances do not necessarily reflect human-perceived temporal alignment qual- ity. Future benchmarks tailored to temporal synchronization and metrics that account for sparse perceptual saliency may bridge this gap. Event Curve Robustness to Non-Semantic Changes. To test the robust- ness of DINOv2 to non-semantic visual changes that may affect event curves, we perform random non-semantic augmentations to 3,450 OES-Pub frames (fps=1), where each frame is randomly perturbed: subpixel translation ±4 px, rotation ±4 ◦ , brightness step [0.4, 1.6]×, or gamma change [0.4, 1.8]. The cosine similar- ity between DINOv2 features before and after augmentation is μ = 0.983 with σ = 0.025, indicating that the video event curves remain robust under these perturbations. Representation Power of Event Curves. Event curves give relative tim- ing information, while text conveys complex semantic guidance. To further measure the representation power of event curves, we set up a 3-way classi- fication task among cinematic/natural/dance videos using samples from OES- Pub/MovieGenBench/AIST++. Using the video-event curves as input, a 1-layer MLP with a 90%/10% train–test split achieved 68.2% test accuracy, showing that the curves contain meaningful information based on video dynamics. This shows that video-event curves already capture non-trivial video characteristics, which are further complemented by text prompts in the full model. D Qualitative Results on Event Curves Event Curves over Diverse Dynamics. V2M-Zero is conditioned on both event curves and text. Event curves with normalization provide relative tem- 20Lin et al. poral structure, aiming to represent when changes happen, while text conveys complementary information such as semantic content, energy, and mood. To validate this behavior, Figure 5 shows two videos with distinct visual dynamics together with an example generated music event curve. The blue solid curve corresponds to a video with frequent scene cuts (i.e., Good Sample 1 in demo, (Sora2) 20251120_0147_01kaengrpqetrrwv7jyycmnkzn.mp4), while the orange curve corresponds to a video dominated by slower camera motion, such as pans (i.e., Random Sample, (MovieGen) 148.mp4, in demo). Although both visual signals are normalized, their temporal patterns remain clearly different: the scene-cut video exhibits sharp, sparse peaks, whereas the slow-pan video varies more smoothly over time. This example shows that normalization does not erase temporal structure; instead, it preserves relative changes that are useful for synchronization. Fine-Grained Temporal Precision. Although event curves are extracted us- ing temporal smoothing, the resulting peaks (i.e., video 1 in Figure 5) still retain precise temporal localization. We use a Hann window, whose central lobe pre- serves peak locations while suppressing high-frequency noise. In equivalent band- width, our smoothing corresponds to roughly 21 samples at 30 fps, or about a 700 ms window. In practice, this smoothing improves robustness without pre- venting accurate localization of salient events, which is consistent with our use of a ±100 ms tolerance in Scene Cut Hit (SCH). V2M-Zero21 References 1. Andrea Agostinelli, Timo I Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, An- toine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, et al. MusicLM: Generating music from text. arXiv Preprint, 2023. 1, 4 2. Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, et al. V-JEPA 2: Self-supervised video models enable understanding, prediction and planning. arXiv Preprint, 2025. 8, 13 3. Yatong Bai, Jonah Casebeer, Somayeh Sojoudi, and Nicholas J. Bryan. DRAGON: Distributional rewards optimize diffusion generative models. TMLR, 2025. 8 4. Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al. AudioLM: a language modeling approach to audio generation. TASLP, 2023. 4 5. Dimitrios Bralios, Jonah Casebeer, and Paris Smaragdis. Re-bottleneck: Latent re-structuring for neural audio autoencoders. In MLSP, 2025. 8 6. Dimitrios Bralios, Paris Smaragdis, and Jonah Casebeer. Learning to upsample and upmix audio in the latent domain. arXiv Preprint, 2025. 8 7. Jonah Casebeer, Ge Zhu, Zhepei Wang, and Nicholas J Bryan. A generative-first neural audio autoencoder. arXiv:2602.15749, 2026. 5, 13 8. Kan Chen, Chuanxi Zhang, Chen Fang, Zhaowen Wang, Trung Bui, and Ram Nevatia. Visually indicated sound generation by perceptually optimized classifi- cation. In ECCVW, 2018. 4 9. Peihao Chen, Yang Zhang, Mingkui Tan, Hongdong Xiao, Deng Huang, and Chuang Gan. Generating visually aligned sound from videos. TIP, 2020. 4 10. Ziyang Chen, Daniel Geng, and Andrew Owens. Images that sound: Composing images and sounds on a single canvas. In NeurIPS, 2024. 4 11. Ziyang Chen, Prem Seetharaman, Bryan Russell, Oriol Nieto, David Bourgin, Andrew Owens, and Justin Salamon. Video-guided foley sound generation with multimodal controls. In CVPR, 2025. 4 12. Ho Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya, Alexander Schwing, and Yuki Mitsufuji. Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis. In CVPR, 2025. 4 13. Sanjoy Chowdhury, Sayan Nag, KJ Joseph, Balaji Vasan Srinivasan, and Dinesh Manocha. MeLFusion: Synthesizing music from image and language cues using diffusion models. In CVPR, 2024. 2, 4 14. Ruben Ciranni, Giorgio Mariani, Michele Mancusi, Emilian Postolache, Giorgio Fabbro, Emanuele Rodolà, and Luca Cosmo. Cocola: Coherence-oriented con- trastive learning of musical audio representations. arXiv Preprint, 2024. 12 15. Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. Simple and controllable music generation. In NeurIPS, 2023. 1, 4 16. Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. High fidelity neural audio compression. arXiv Preprint, 2022. 4 17. Shangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang, Leyan Zhu, Zexin He, Hongming Liu, and Shuicheng Yan. Video background music generation with controllable music transformer. In ACM M, 2021. 4, 12 18. Yuexi Du, Ziyang Chen, Justin Salamon, Bryan Russell, and Andrew Owens. Conditional generation of audio from video via foley analogies. In CVPR, 2023. 4 22Lin et al. 19. Zach Evans, CJ Carr, Josiah Taylor, Scott H Hawley, and Jordi Pons. Fast timing- conditioned latent audio diffusion. arXiv Preprint, 2024. 4 20. Zach Evans, Julian D Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. Stable audio open. In ICASSP, 2025. 4, 6, 8 21. Zhengcong Fei, Mingyuan Fan, Changqian Yu, and Junshi Huang. Flux that plays music. arXiv Preprint, 2024. 1, 4, 6 22. Jonathan Foote and Matthew Cooper. Visualizing musical structure and rhythm via self-similarity. In ICMC, 2001. 6 23. Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya Poria. Text-to-audio generation using instruction guided latent diffusion model. In ACM M, 2023. 4, 6 24. Junmin Gong, Sean Zhao, Sen Wang, Shengyuan Xu, and Joe Guo. ACE-Step: A step towards music generation foundation model. arXiv Preprint, 2025. 1, 4 25. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models. arXiv Preprint, 2024. 14 26. Noor Hammad, C Ailie Fraser, Erik Harpstead, Jessica Hammer, and Mira Dontcheva. “it’s more of a vibe i’m going for”: Designing text-to-music gener- ation interfaces for video creators. In DIS, 2025. 2, 3, 5, 8, 14, 16 27. Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al. Cnn architectures for large-scale audio classification. In ICASSP, 2017. 9 28. Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv Preprint, 2022. 6 29. Sture Holm. A simple sequentially rejective multiple test procedure. Scandinavian journal of statistics, 1979. 11 30. Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren, Luping Liu, Mingze Li, Zhenhui Ye, Jinglin Liu, Xiang Yin, and Zhou Zhao. Make-an-audio: Text-to- audio generation with prompt-enhanced diffusion models. In ICML, 2023. 4 31. Fathinah Izzati, Xinyue Li, Yuxuan Wu, and Gus Xia. MusiScene: Leveraging mu-llama for scene imagination and enhanced video background music generation. arXiv Preprint, 2025. 4 32. Jaeyong Kang, Soujanya Poria, and Dorien Herremans. Video2music: Suitable music generation from videos using an affective multimodal transformer model. arXiv Preprint, 2023. 4 33. Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Cotracker: It is better to track together. In ECCV, 2024. 8, 12, 14 34. Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi. Fr\’echet audio distance: A metric for evaluating music enhancement algorithms. arXiv Preprint, 2018. 9 35. Haven Kim, Zachary Novack, Weihan Xu, Julian McAuley, and Hao-Wen Dong. Video-guided text-to-music generation using public domain movie collections. In ISMIR, 2025. 2, 4, 8, 10 36. Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. Audiogen: Textually guided audio generation. In ICLR, 2023. 4 37. Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kun- dan Kumar. High-fidelity audio compression with improved RVQGAN. In NeurIPS, 2023. 4 V2M-Zero23 38. Saksham Singh Kushwaha and Yapeng Tian. VinTAGe: Joint video and text conditioning for holistic audio generation. In CVPR, 2025. 4 39. Heeseung Kwon, Manjin Kim, Suha Kwak, and Minsu Cho. Learning self- similarity in space and time as generalized motion for video action recognition. In ICCV, 2021. 6 40. Max WY Lam, Qiao Tian, Tang Li, Zongyu Yin, Siyuan Feng, Ming Tu, Yuliang Ji, Rui Xia, Mingbo Ma, Xuchen Song, et al. Efficient neural music generation. In NeurIPS, 2023. 1, 4 41. Hsin-Ying Lee, Xiaodong Yang, Ming-Yu Liu, Ting-Chun Wang, Yu-Ding Lu, Ming-Hsuan Yang, and Jan Kautz. Dancing to music. In NeurIPS, 2019. 4 42. Jiajun Li, Tianze Xu, Xuesong Chen, Xinrui Yao, Jingchou Han, and Shuchang Liu. Mozart’s touch: a lightweight multimodal music generation framework based on pre-trained large models. In AIGC, 2025. 4 43. Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa. AI choreographer: Music conditioned 3d dance generation with AIST++. In ICCV, 2021. 9, 11, 12 44. Ruiqi Li, Siqi Zheng, Xize Cheng, Ziang Zhang, Shengpeng Ji, and Zhou Zhao. MuVi: Video-to-music generation with semantic alignment and rhythmic synchro- nization. arXiv Preprint, 2024. 2, 4 45. Sifei Li, Weiming Dong, Yuxin Zhang, Fan Tang, Chongyang Ma, Oliver Deussen, Tong-Yee Lee, and Changsheng Xu. Dance-to-music generation with encoder- based textual inversion. In SIGGRAPH Asia, 2024. 4, 9, 12, 16 46. Sizhe Li, Yiming Qin, Minghang Zheng, Xin Jin, and Yang Liu. Diff-BGM: A diffusion model for video background music generation. In CVPR, 2024. 4 47. Sifei Li, Binxin Yang, Chunji Yin, Chong Sun, Yuxin Zhang, Weiming Dong, and Chen Li. VidMusician: Video-to-music generation with semantic-rhythmic alignment via hierarchical visual features. arXiv Preprint, 2024. 2, 4 48. Yan-Bo Lin and Gedas Bertasius. Siamese vision transformers are scalable audio- visual learners. In ECCV, 2024. 8, 13 49. Yan-Bo Lin, Yu Tian, Linjie Yang, Gedas Bertasius, and Heng Wang. VMAS: Video-to-music generation via semantic alignment in web music videos. In WACV, 2025. 2, 4, 9 50. Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley. AudioLDM: Text-to-audio generation with latent diffusion models. In ICML, 2023. 4, 6 51. Haohe Liu, Qiao Tian, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D Plumbley. AudioLDM 2: Learn- ing holistic audio generation with self-supervised pretraining. arXiv Preprint, 2023. 4, 6 52. Huadai Liu, Jialei Wang, Kaicheng Luo, Wen Wang, Qian Chen, Zhou Zhao, and Wei Xue. ThinkSound: Chain-of-thought reasoning in multimodal large language models for audio generation and editing. In NeurIPS, 2025. 4 53. Shansong Liu, Atin Sakkeer Hussain, Qilong Wu, Chenshuo Sun, and Ying Shan. M 2 UGen: Multi-modal music understanding and generation with the power of large language models. arXiv Preprint, 2023. 4 54. Shansong Liu, Atin Sakkeer Hussain, Qilong Wu, Chenshuo Sun, and Ying Shan. MuMu-LLaMA: Multi-modal music understanding and generation via large lan- guage models. arXiv Preprint, 2024. 9, 10, 18 55. Xiulong Liu, Kun Su, and Eli Shlizerman. Tell what you hear from what you see-video to audio generation through text. In NeurIPS, 2024. 4 56. Xiaohao Liu, Teng Tu, Yunshan Ma, and Tat-Seng Chua. Extending visual dy- namics for video-to-music generation. arXiv Preprint, 2025. 2, 4 24Lin et al. 57. Zihan Liu, Shuangrui Ding, Zhixiong Zhang, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Dahua Lin, and Jiaqi Wang. SongGen: A single stage auto- regressive transformer for text-to-song generation. In ICML, 2025. 1, 4 58. Simian Luo, Chuanhao Yan, Chenxu Hu, and Hang Zhao. Diff-Foley: Synchro- nized video-to-audio synthesis with latent diffusion models. In NeurIPS, 2023. 4 59. Ilaria Manco, Benno Weck, Seungheon Doh, Minz Won, Yixiao Zhang, Dmitry Bogdanov, Yusong Wu, Ke Chen, Philip Tovstogan, Emmanouil Benetos, et al. The song describer dataset: a corpus of audio captions for music-and-language evaluation. arXiv Preprint, 2023. 9, 10 60. Xinhao Mei, Varun Nagaraja, Gael Le Lan, Zhaoheng Ni, Ernie Chang, Yangyang Shi, and Vikas Chandra. FoleyGen: Visually-guided audio generation. arXiv Preprint, 2023. 4 61. Jan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, and Soujanya Poria. Mustango: Toward controllable text-to-music generation. arXiv Preprint, 2023. 4, 6 62. Zachary Novack, Zach Evans, Zack Zukowski, Josiah Taylor, CJ Carr, Julian Parker, Adnan Al-Sinan, Gian Marco Iodice, Julian McAuley, Taylor Berg- Kirkpatrick, et al. Fast text-to-audio generation with adversarial post-training. arXiv Preprint, 2025. 4 63. Zachary Novack, Julian McAuley, Taylor Berg-Kirkpatrick, and Nicholas J. Bryan. DITTO: Diffusion inference-time t-optimization for music generation. In ICML, 2025. 1, 4 64. Zachary Novack, Ge Zhu, Jonah Casebeer, Julian McAuley, Taylor Berg- Kirkpatrick, and Nicholas J. Bryan. Presto! distilling steps and layers for ac- celerating music generation. In ICLR, 2025. 8 65. Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El- Nouby, et al. DINOv2: Learning robust visual features without supervision. arXiv Preprint, 2023. 7, 8, 10, 13, 18 66. Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman. Visually indicated sounds. In CVPR, 2016. 4 67. William Peebles and Saining Xie. Scalable diffusion models with transformers. In ICCV, pages 4195–4205, 2023. 6 68. Geoffroy Peeters. Self-similarity-based and novelty-based loss for music structure analysis. In International Society of Music Information Retreival, 2023. 6 69. Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, et al. Movie Gen: A cast of media foundation models. arXiv Preprint, 2024. 9, 10 70. Fan Qi, Kunsheng Ma, and Changsheng Xu. Customized condition controllable generation for video soundtrack. In CVPR, 2025. 4 71. Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision. In ICML, 2023. 16 72. Ciara Rowles, Varun Jampani, Simon Donné, Shimon Vainer, Julian Parker, and Zach Evans. Foley control: Aligning a frozen latent text-to-audio model to video. arXiv Preprint, 2025. 4 73. Flavio Schneider, Ojasv Kamal, Zhijing Jin, and Bernhard Schölkopf. Moûsai: Text-to-music generation with long-context latent diffusion. arXiv Preprint, 2023. 1, 4 74. Sizhe Shan, Qiulin Li, Yutao Cui, Miles Yang, Yuehai Wang, Qun Yang, Jin Zhou, and Zhao Zhong. Hunyuanvideo-foley: Multimodal diffusion with representation alignment for high-fidelity foley audio generation. arXiv Preprint, 2025. 17 V2M-Zero25 75. Megha Sharma, Muhammad Taimoor Haseeb, Gus Xia, and Yoshimasa Tsuruoka. M2M-Gen: A multimodal framework for automated background music generation in japanese manga using large language models. arXiv Preprint, 2024. 4 76. Eli Shechtman and Michal Irani. Matching local self-similarities across images and videos. In CVPR, 2007. 6 77. Eli Shlizerman, Lucio Dery, Hayden Schoen, and Ira Kemelmacher-Shlizerman. Audio to body dynamics. In CVPR, 2018. 4 78. Kun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin, Joonseok Lee, Chris Donahue, Fei Sha, Aren Jansen, Yu Wang, Mauro Verzetti, et al. V2Meow: Me- owing to the visual beat via music generation. In AAAI, 2024. 4 79. Kun Su, Xiulong Liu, and Eli Shlizerman. From vision to audio and beyond: A unified model for audio-visual representation and generation. In ICML, 2024. 4 80. Changchang Sun, Gaowen Liu, Charles Fleming, and Yan Yan. Enhancing dance- to-music generation via negative conditioning latent diffusion model. In CVPR, 2025. 4 81. Or Tal, Felix Kreuk, and Yossi Adi. Auto-regressive vs flow-matching: a compar- ative study of modeling paradigms for text-to-music generation. TMLR, 2025. 1, 4 82. Vanessa Tan, Junghyun Nam, Juhan Nam, and Junyong Noh. Motion to dance music generation using latent diffusion model. In SIGGRAPH Asia, 2023. 12 83. Zineng Tang, Ziyi Yang, Mahmoud Khademi, Yang Liu, Chenguang Zhu, and Mo- hit Bansal. Codi-2: In-context, interleaved, and interactive any-to-any generation. arXiv Preprint, 2023. 4 84. Zineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng, and Mohit Bansal. Any- to-any generation via composable diffusion. In NeurIPS, 2023. 4 85. Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, et al. Gemma 3 technical report. arXiv Preprint, 2025. 8, 14, 16 86. Sida Tian, Can Zhang, Wei Yuan, Wei Tan, and Wenjie Zhu. XMusic: Towards a generalized and controllable symbolic music generation framework. TMM, 2025. 4 87. Zeyue Tian, Yizhu Jin, Zhaoyang Liu, Ruibin Yuan, Xu Tan, Qifeng Chen, Wei Xue, and Yike Guo. Audiox: Diffusion transformer for anything-to-audio gener- ation. arXiv Preprint, 2025. 2, 4, 9, 10, 18 88. Zeyue Tian, Zhaoyang Liu, Ruibin Yuan, Jiahao Pan, Qifeng Liu, Xu Tan, Qifeng Chen, Wei Xue, and Yike Guo. VidMuse: A simple video-to-music generation framework with long-short-term modeling. In CVPR, 2025. 2, 4, 9, 10, 18 89. Xinyi Tong, Yiran Zh, Jishang Chen, Chunru Zhan, Tianle Wang, Sirui Zhang, Nian Liu, Tiezheng Ge, Duo Xu, Xin Jin, Feng Yu, and Song-Chun Zhu. Video echoed in music: Semantic, temporal, and rhythmic alignment for video-to-music generation. arXiv Preprint, 2025. 4 90. Shuhei Tsuchida, Satoru Fukayama, Masahiro Hamasaki, and Masataka Goto. AIST dance video database: Multi-genre, multi-dancer, and multi-camera database for dance information processing. In ISMIR, 2019. 9, 11, 12 91. Apoorv Vyas, Bowen Shi, Matthew Le, Andros Tjandra, Yi-Chiao Wu, Baishan Guo, Jiemin Zhang, Xinyue Zhang, Robert Adkins, William Ngan, et al. Au- diobox: Unified audio generation with natural language prompts. arXiv Preprint, 2023. 4 92. Baisen Wang, Le Zhuo, Zhaokai Wang, Chenxi Bao, Wu Chengjing, Xuecheng Nie, Jiao Dai, Jizhong Han, Yue Liao, and Si Liu. Multimodal music generation with explicit bridges and retrieval augmentation. arXiv Preprint, 2024. 4 26Lin et al. 93. Jun Wang, Xijuan Zeng, Chunyu Qiang, Ruilong Chen, Shiyao Wang, Le Wang, Wangjing Zhou, Pengfei Cai, Jiahui Zhao, Nan Li, et al. Kling-Foley: Multimodal diffusion transformer for high-quality video-to-audio generation. arXiv Preprint, 2025. 4 94. Minz Won, Yun-Ning Hung, and Duc Le. A foundation model for music infor- matics. In ICASSP, 2024. 7, 8, 13, 18 95. Junxian Wu, Weitao You, Heda Zuo, Dengming Zhang, Pei Chen, and Lingyun Sun. Controllable video-to-music generation with multiple time-varying condi- tions. In ACM M, 2025. 2, 4, 9, 10, 18 96. Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation. In ICASSP, 2023. 9 97. Zhifeng Xie, Qile He, Youjia Zhu, Qiwei He, and Mengtian Li. FilmComposer: Llm-driven music production for silent film clips. In CVPR, 2025. 2, 4 98. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv Preprint, 2025. 14 99. Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu. Diffsound: Discrete diffusion model for text-to-sound generation. TASLP, 2023. 4 100. Jiashuo Yu, Yaohui Wang, Xinyuan Chen, Xiao Sun, and Yu Qiao. Long-term rhythmic video soundtracker. In ICML, 2023. 4, 12 101. Ruibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang, Jiahao Pan, Yongyi Zang, Haohe Liu, Yiming Liang, Wenye Ma, Xingjian Du, et al. Yue: Scaling open foundation models for long-form music generation. arXiv Preprint, 2025. 1, 4 102. Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. Soundstream: An end-to-end neural audio codec. TASLP, 2021. 4 103. Liqian Zhang and Magdalena Fuentes. SONIQUE: Video background music gen- eration using unpaired audio-visual data. In ICASSP, 2025. 2, 4, 9, 10, 18 104. Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui, and Tamara L Berg. Visual to sound: Generating natural sound for videos in the wild. In CVPR, 2018. 4 105. Zitang Zhou, Ke Mei, Yu Lu, Tianyi Wang, and Fengyun Rao. Harmonyset: A comprehensive dataset for understanding video-music semantic alignment and temporal synchronization. In CVPR, 2025. 4 106. Ye Zhu, Kyle Olszewski, Yu Wu, Panos Achlioptas, Menglei Chai, Yan Yan, and Sergey Tulyakov. Quantized GAN for complex music generation from dance videos. In ECCV, 2022. 4 107. Ye Zhu, Yu Wu, Kyle Olszewski, Jian Ren, Sergey Tulyakov, and Yan Yan. Dis- crete contrastive diffusion for cross-modal music and image generation. In ICLR, 2023. 4, 12 108. Wenlin Zhuang, Congyi Wang, Jinxiang Chai, Yangang Wang, Ming Shao, and Siyu Xia. Music2dance: Dancenet for music-driven dance generation. TOMM, 2022. 4 109. Le Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao, Chenxi Bao, Stanley Peng, Song- hao Han, Aixi Zhang, Fei Fang, and Si Liu. Video background music generation: Dataset, method and evaluation. In ICCV, 2023. 4 110. Alon Ziv, Itai Gat, Gael Le Lan, Tal Remez, Felix Kreuk, Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. Masked audio generation using a single non-autoregressive transformer. In ICLR, 2024. 1, 4 111. Heda Zuo, Weitao You, Junxian Wu, Shihong Ren, Pei Chen, Mingxu Zhou, Yujia Lu, and Lingyun Sun. GVMGen: A general video-to-music generation model with hierarchical attentions. In AAAI, 2025. 2, 4, 9, 10, 18