Paper deep dive
Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping
Jingwei Zhao, Gus Xia, Ziyu Wang, Ye Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/5/2026, 4:56:34 AM
Summary
This paper introduces a cross-modal framework for learning implicit music styles from raw audio to condition symbolic music generation for piano arrangements. The model utilizes a Querying Transformer (Q-Former) to extract style representations from a pre-trained audio language model and applies them to condition a symbolic language model. The training involves a two-stage strategy: contrastive learning to align auditory style with symbolic expression, followed by generative modeling. The approach enables controllable piano cover generation, style transfer, and audio-to-MIDI retrieval, demonstrating superior style-aware alignment and music quality compared to existing baselines.
Entities (15)
Relation Signals (10)
Q-Former → extracts → Style Representation
confidence 95% · our model leverages a Querying Transformer (Q-Former) to extract style representations from a large, pre-trained audio language model
PIAST → usedfortraining → Our Model
confidence 95% · We use POP909 [36] and PIAST for training.
POP909 → usedfortraining → Our Model
confidence 95% · We use POP909 [36] and PIAST for training.
Q-Former → trainedwith → Contrastive Learning
confidence 94% · The first stage employs contrastive learning, training the Q-Former to extract auditory representations
Q-Former → conditions → Symbolic Language Model
confidence 93% · applies them to condition a symbolic LM for generating piano arrangements
MusicGen → usedasbackbonefor → Q-Former
confidence 92% · We integrate the Q-Former into MusicGen [24], one of the leading music audio LMs available today.
MuseCoco → usedasbackbonefor → Symbolic Generation
confidence 92% · we take advantage of the generative capability of MuseCoco [3], a large-scale symbolic music LM.
Q-Former → bridges → Audio Modality
confidence 90% · A Q-Former module bridges the modality gap between a frozen audio LM and a symbolic music LM.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit and hidden in concrete music examples. In this paper, we introduce a cross-modal framework that learns implicit music styles from raw audio and applies them to symbolic music generation. Inspired by BLIP-2, our model leverages a Querying Transformer (Q-Former) to extract style representations from a large, pre-trained audio language model (LM), and further applies them to condition a symbolic LM for generating piano arrangements. We adopt a two-stage training strategy: contrastive learning to align auditory style with symbolic expression, followed by generative modeling for music arrangement. Our model generates piano performances jointly conditioned on a lead sheet (content) and a reference audio example (style), enabling controllable and stylistically faithful arrangement. Experiments demonstrate the effectiveness of our approach in piano cover generation, style transfer, and audio-to-MIDI retrieval, achieving substantial improvements in style-aware alignment and music quality.
Tags
Links
- Source: https://arxiv.org/abs/2608.03050v1
- Canonical: https://arxiv.org/abs/2608.03050v1
Trouble viewing inline? Open PDF directly →
Full Text
62,021 characters extracted from source content.
Expand or collapse full text
LEARNING MUSIC STYLE FOR PIANO ARRANGEMENT THROUGH CROSS-MODAL BOOTSTRAPPING Jingwei Zhao 1,4 Gus Xia 2 Ziyu Wang 2,3 Ye Wang 4 1 Songscription 2 Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) 3 Courant Institute, New York University 4 School of Computing, National University of Singapore jzhao@u.nus.edu, gus.xia@mbzuai.ac.ae, ziyu.wang@nyu.edu, dcswangy@nus.edu.sg ABSTRACT What is music style? Though often described using text labels such as “swing,” “classical,” or “emotional,” the real style remains implicit and hidden in concrete music exam- ples. In this paper, we introduce a cross-modal framework that learns implicit music styles from raw audio and applies them to symbolic music generation. Inspired by BLIP-2, our model leverages a Querying Transformer (Q-Former) to extract style representations from a large, pre-trained audio language model (LM), and further applies them to condition a symbolic LM for generating piano arrange- ments. We adopt a two-stage training strategy: contrastive learning to align auditory style with symbolic expression, followed by generative modeling for music arrangement. Our model generates piano performances jointly condi- tioned on a lead sheet (content) and a reference audio ex- ample (style), enabling controllable and stylistically faith- ful arrangement. Experiments demonstrate the effective- ness of our approach in piano cover generation, style trans- fer, and audio-to-MIDI retrieval, achieving substantial im- provements in style-aware alignment and music quality. 1 1. INTRODUCTION Automatic music generation is often controlled by explicit content such as melody, chords, and text labels [1–4], but music concepts can be more nuanced than we often real- ize. When musicians learn a style, instead of relying on abstract descriptors like “romantic” or “jazz” alone, they absorb patterns from music examples that share common stylistic traits. The commonality across these examples forms a style, an implicit one that cannot be fully described with words or labels but only understood through the music itself. This paper studies such implicit style qualities, par- ticularly regarding grooving patterns and dynamics in pi- ano performances, where an abstract descriptor often over- simplifies their richness even though they are immediately 1 Demo page: https://zhaojw1998.github.io/bossa/ © J. Zhao, G. Xia, Z. Wang, and Y. Wang. Licensed under a Creative Commons Attribution 4.0 International License (C BY 4.0). Attribution: J. Zhao, G. Xia, Z. Wang, and Y. Wang, “Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping”, in Proc. of the 27th Int. Society for Music Information Retrieval Conf., Abu Dhabi, UAE, 2026. Symbolic LM Audio LM Audio Q-Former Lead Sheet ⋯ Piano Arrangement Style Representation (providing music content) (in ragtime) (e.g., a ragtime song) Stage 1 Cross-Modal Representation Learning Stage 2 Generative Modeling Figure 1: A Q-Former module bridges the modality gap between a frozen audio LM and a symbolic music LM. It extracts cross-modal music style from the hidden repre- sentations of the audio LM and, together with a lead sheet providing music content, conditions the symbolic LM for piano arrangement. The Q-Former is trained in a two- stage process, effectively bootstrapping audio-to-symbolic arrangement without re-training either LM backbone. perceivable in audio. We explore how such implicit style can be internalized from given audio examples and control music generation in a deep learning framework. Large-scale music language models (music LMs) have shown strong capabilities in learning explicit music con- tent, as demonstrated by probing studies [5–9] and adapter- based designs [10–13]. Yet, control over implicit style re- mains limited. For example, in audio-to-MIDI generation, existing models can extract melody and chords [14,15], but capturing stylistic traits like the rhythmic feel and more expressive nuances remains a greater challenge. This re- quires disentangling style from music content, which cur- rent music LM-based studies have yet to explore. In this paper, we explore learning implicit music style in a cross-modal setting for symbolic piano arrangement. Our goal is to generate an expressive piano MIDI performance conditioned on two inputs: an audio example (providing style) and a lead sheet MIDI score (melody and chords as content). To achieve this, we connect pre-trained music LMs in the audio and symbolic domains using a Querying Transformer (Q-Former), a lightweight Transformer orig- inally designed for vision-language alignment [16]. As shown in Figure 1, we extend the Q-Former to capture a style representation from the hidden states of an audio LM. Then, a symbolic LM conditions on the style representa- tion, along with the content of the lead sheet, to generate arXiv:2608.03050v1 [cs.SD] 4 Aug 2026 a piano arrangement. The Q-Former enables cross-modal style transfer between two large unimodal LMs without re- training them—a process we refer to as bootstrapping. In our design, we treat the Q-Former as a bottleneck to transfer only style-related information and adopt a two- stage training strategy. The first stage employs contrastive learning, training the Q-Former to extract auditory repre- sentations that are musically relevant, expressible in sym- bolic piano arrangements, and independent of explicit mu- sic content. These include accompaniment textures, groov- ing patterns, and performance dynamics (MIDI velocity contour and tempo). The second stage focuses on gener- ative modeling, where the Q-Former’s output conditions the symbolic LM to generate a piano arrangement with the desired expression. We show that the complete sys- tem generates more stylistically accurate cover songs com- pared to existing audio-to-symbolic arrangement methods, while also enabling piano style transfer by conditioning on alternative audio examples. In addition, we demon- strate that the Q-Former can also be applied to audio-to- symbolic retrieval, further highlighting its strength as a general-purpose cross-modal representation learner. In sum, the contributions of this paper are threefold: 1. We use Q-Former to align audio and symbolic modalities via implicit music style, extending its role beyond content alignment in vision-language tasks. 2. We present a new methodology to disentangle music style from large, pre-trained LMs, offering a more scalable alternative to traditional latent-variable dis- entanglement methods. 3. Our model achieves style-aware audio-to-symbolic piano cover arrangement. Experiments demonstrate that it outperforms existing audio-to-symbolic mod- els, including both disentanglement-based methods and standard LM approaches. 2. RELATED WORK We review two relevant areas. Section 2.1 overviews re- cent advances in music LMs, while Section 2.2 focuses on piano cover generation, a primary task of this paper. 2.1 Music Language Models Rapid progress in large-scale language models has trans- formed how we interact with various forms of media, in- cluding text, image, and music [16–20]. In particular, large music LMs [4, 21–23] have notably influenced cre- ative practices and user experiences. Models like Mu- sicGen [24] generates music audio with rich timbres di- rectly from text, while MuseCoco [3] produces symbolic compositions with well-structured textures in varied gen- res. These advancements are driven by training large-scale neural networks on extensive data, scaling up to billions of parameters to enhance controllability and musicality. Despite these successes, most existing music LMs op- erate in a unimodal setting, focusing solely on either au- dio or symbolic representations. Although text-to-music generation has been increasingly effective [4, 21, 22], text descriptions may fall short in expressing nuanced style or performance subtlety. In contrast, our work explores a cross-modal framework that bridges audio and symbolic modalities. This approach enables more intuitive and fine- grained control over music style beyond what can be con- veyed through text alone. 2.2 Piano Cover Generation Piano cover generation aims to reinterpret an audio record- ing as a symbolic piano performance. Unlike traditional music transcription, which primarily analyzes note-level content such as pitch and timing [25–29], a piano cover often targets higher-level, more structured music elements that shape the feel of a performance. The goal is to gener- ate symbolic arrangements that not only sound correct but also feel musically aligned with the original audio. Existing approaches to piano cover generation often leverage pre-trained transcription models, which primar- ily extract melodic and harmonic content from the au- dio [15, 30–33]. However, such models tend to overlook stylistic nuances, resulting in outputs accurate in harmony but lacking the expressive character of the source perfor- mance. In this paper, we re-frame piano cover generation through the lens of content-style disentanglement, acquir- ing content in the symbolic form (i.e., melody and chord progression) while learning style from the audio. This ap- proach bridges the audio-symbolic gap more effectively, capturing not just what is played, but how it is played. 3. METHOD To bridge the modality gap from audio to symbolic mu- sic, we adopt the Q-Former [16] under a two-stage training strategy, as shown in Figure 2. In Section 3.1, we first introduce our audio-symbolic data pairing method that fa- cilitates style learning. We illustrate the Q-Former archi- tecture in Section 3.2, followed by the two-stage training procedure in Sections 3.3 and 3.4. We provide more model configuration details in Appendix A. 3.1 Data Pairing for Style Learning Training an audio-to-symbolic alignment model requires paired audio–MIDI data. In this work, we use 10s audio clips paired with 4-bar MIDI segments. Since our goal is to capture style rather than low-level note transcription, we construct the pairs to be loosely aligned. Specifically, for each audio clip, we select a MIDI segment near its center with a random temporal shift of up to±1 second, and we randomly transpose the MIDI into all 12 keys. This design assumes that music style is locally consistent, while dis- couraging the model from memorizing exact note-to-note correspondences. In the following sections, we show that the two-step training method enables the model to abstract style features that are shared between the modalities. We represent music audio as raw waveforms sampled at 32kHz. MIDI is tokenized into note event sequences quan- tized at 1/12-beat resolution. We include various symbolic Object 3: Generative Object 1: Contrastive Object 2: Matching Querying Embeddings Q-Former ⋯ Audio LM ⋯ Symbolic Note Tokens Self-AttentionSelf-Attention Cross-Attention Feedforward Feedforward Mask Control ×24 퐙∈ℝ 퐾×768 Audio (Randomly Initialized) (a) Stage-I: Cross-modal learning with Q-Former. ⋯ 퐙 (from Q-Former) Lead Sheet Tokens Output Piano Arrangement Tokens Piano Arrangement Tokens Symbolic Music LM LoRA Linear (shift-right) (b) Stage-I: Audio-to-symbolic arrangement. Figure 2: Overview of the two-stage framework. (Left) Stage-I: The Q-Former is a Transformer encoder with two parallel, modality-specific streams (audio and symbolic). Both streams are used during training for cross-modal alignment, while only the audio stream is retained at test time. It extracts a cross-modal music style representationZ via cross-attention to an audio LM. (Right) Stage-I: A symbolic music LM further generates piano arrangements conditioned on (i) the style embeddingZ obtained from the Q-Former and (i) a lead sheet providing music content. features including time signature (quadruple and triple me- ters), tempo curve, note pitch, duration, and velocity. 3.2 Q-Former Architecture The Q-Former is a Transformer encoder that processes two parallel input streams (audio and symbolic modalities) and learns a shared cross-modal music style representationZ. As shown in Figure 2a, the left stream is connected to the audio LM via cross-attention, while the right stream en- codes symbolic piano arrangement tokens via the shared self-attention. A set of K query embeddings (queries), ini- tialized randomly, is fed into the left stream to extract style cues from the audio modality while being aligned with the symbolic modality. At test time, only the left stream is retained to uncover cross-modal style directly from audio. 3.3 Stage-I: Audio-Symbolic Representation Learning We integrate the Q-Former into MusicGen [24], one of the leading music audio LMs available today. The queries in- teract with MusicGen’s audio hidden states through cross- attention and remain connected to the symbolic stream via the shared self-attention layers. To encourage cross-modal style abstraction, we introduce three complementary train- ing objectives, each paired with a tailored self-attention mask that regulates cross-modal interactions, as illustrated in Figure 2a and detailed below. The primary objective is Audio-Symbolic Contrastive Learning, which enforces a higher audio-symbolic sim- ilarity for positive (original) pairs compared to negative ones (i.e., randomly paired audio and MIDI clips). Let Z∈R K×768 be the K query outputs from the audio stream of Q-Former, and t ∈R 1×768 be the output embedding of the start token (<s>) from the symbolic stream. We de- fine the audio-symbolic similarity as max k (cos(Z k , t)) for k = 1, 2,· , K, where cos(·,·) denotes the cosine simi- larity. The contrastive loss pulls closer aligned audio and symbolic clips in the representation space, while pushing apart unrelated pairs. To prevent information leakage, we employ a unimodal self-attention mask, ensuring queries and symbolic tokens do not attend to each other. For de- tailed mask design, we refer readers to BLIP-2 [16]. The second objective is Audio-Symbolic Matching. It is formulated as a binary classification task, where the model predicts whether a given audio-symbolic pair cor- responds to each other. On top of contrastive loss, the matching loss aims to capture a finer cross-modal corre- spondence. In this case, we apply no masking, allowing the queries to attend across modalities. Each query output Z k is fed into a binary linear classifier to produce a logit, and the logits from all queries are averaged to compute the final matching score. To create informative negative pairs, we employ the hard negative mining strategy in [16, 34]. The final objective, Audio-Grounded Symbolic Gen- eration, enforces the Q-Former’s symbolic stream to auto- regressively reconstruct the input piano arrangement. We implement a cross-modal causal self-attention mask, al- lowing the symbolic tokens to attend to the queries but not vice versa. This objective ensures that the style cues ex- tracted from the audio are sufficiently informative to sup- port symbolic realization. To signal a decoding task, we replace the starting <s> token with a special <DEC> to- ken. Additionally, we prepend a sequence of lead sheet to- kens before <DEC> so that the queries are encouraged to extract style, rather than transcribing content from audio. 3.4 Stage-I: Audio-to-Symbolic Generative Modeling In the generative modeling stage, we take advantage of the generative capability of MuseCoco [3], a large-scale sym- bolic music LM. As illustrated in Figure 2b, MuseCoco is used to reconstruct a piano arrangement based on two concatenated conditional inputs: 1) the query output em- beddingsZ from the Q-Former, and 2) a lead sheet. The Q-Former is pre-trained at Stage-I to extract cross-modal music style from the audio, thus providing style guidance. The lead sheet defines the theme melody and chord pro- gression as the content. Since MuseCoco does not natively support lead sheet conditioning, and the inclusion of the lead sheet tokens alters its input format, we insert a LoRA Table 1: Objective evaluation on content preservation and style coherence. Ballroom/GTZAN are out-of-distribution sets to assess generalization to unseen genres and instrumentation. Results are reported in the form of mean ± sem s (in percentage), where sem is the standard error of the mean. Different superscript letters s within a column indicate significant differences (p-value p < 0.05/6) based on Wilcoxon signed-rank test with Bonferroni correction. POP909 (In-Distribution) Test SetBallroom/GTZAN MCA↑CA↑GPC↑VCC↑TA↑MCA↑CA↑TA↑ Ours32.0±0.5 b 33.3±1.1 b 79.2±0.5 a 76.6±0.5 a 83.6±1.5 a 17.8±0.3 a 16.3±0.4 a 79.4±1.1 a PCG239.3±0.6 a 40.0±0.9 a 78.4±0.6 a 73.2±0.5 b 74.4±2.0 c 17.2±0.3 a 15.5±0.5 b 57.6±1.5 c A2M15.0±0.3 c 22.4±0.9 d 69.0±0.7 b 68.6±0.8 c 77.8±1.4 b 10.9±0.2 b 14.8±0.5 b - w/o PT30.0±0.6 b 28.8±1.0 c 78.6±0.5 a 76.3±0.5 a 68.6±1.9 d 17.2±0.3 a 15.0±0.4 b 74.5±1.3 b adapter [35] into each self-attention layer. This enables the model to reweight attention and incorporate the new con- ditioning inputs, while keeping MuseCoco itself frozen. 4. EXPERIMENTS Our model generates piano performances jointly condi- tioned on a lead sheet and an audio reference. When the two inputs are aligned with each other, the task corre- sponds to piano cover generation; when they are unpaired, the task becomes cross-modal style transfer. This section focuses on piano cover generation, which allows direct comparison with prior baselines. The other scenario is cov- ered in Section 5. We introduce the datasets in Section 4.1 and baseline models in Section 4.2. Our evaluation is di- vided into two parts: objective evaluation in Section 4.3, and subjective evaluation in Section 4.4. 4.1 Datasets We use POP909 [36] and PIAST [37] for training. POP909 has 909 piano cover arrangements. The music genre is pri- marily Mandarin pop, and the accompanying audio fea- tures band instrumentation, which can help the model learn generalizable audio representations of pop music. PIAST has 8K piano recordings along with transcriptions across a variety of genres. Despite the lack of band instrumentation in the piano recording, the genre diversity of PIAST en- courages the model to produce more expressive and stylis- tically varied performances. We split both datasets at song level into training (90%), validation (5%), and test (5%) sets. Each symbolic MIDI file is clipped into 4-bar seg- ments with a 2-bar hop size, center-aligned with the cor- responding 10s audio clip. At inference time, longer se- quences are produced via windowed sampling, advancing by 2 bars while conditioning on the preceding 2 bars. We also test on two out-of-distribution datasets: Ball- room [38, 39] and GTZAN [40, 41], both featuring diverse band/orchestral instrumentation and fine-grained music genres such as jive and bossa nova. Since they lack paired symbolic annotations, we use them for testing only. This allows us to assess the model’s generalization ability and its capacity to accommodate styles beyond pop music. 4.2 Baseline Models We compare our model against two representative piano cover generation models: PiCoGen2 [32] and Audio-to- MIDI [15], as well as one ablation variant of our method. PiCoGen2 (PCG2) is a Transformer-based language model that builds on the hidden states of Sheetsage [14,42], which itself is derived from Jukebox [43], a pre-trained, large-scale music language model. Leveraging Jukebox’s internalized understanding of music content, PiCoGen2 generates symbolic piano covers directly from audio. Audio-to-MIDI (A2M) is an auto-encoder-based disen- tanglement framework, using separate modules to extract beat and tempo [44], chord [45], and piano texture from audio. The texture extractor is initialized from a piano tran- scription model [46]. The extracted components are then merged to form a symbolic piano arrangement. Ours w/o Pre-Training (w/o PT) is an ablation vari- ant of our model in which the Q-Former is trained directly in Stage-I, without undergoing the representation learning phase in Stage-I. This setup tests the validity of the two- stage training strategy we applied in this work. To ensure a fair comparison with baseline models, we use Sheetsage [14, 42] to transcribe lead sheets from the audio, making audio the sole input for all methods. 4.3 Objective Evaluation Piano cover generation transforms an audio input into a symbolic piano performance, which should ideally cap- ture not only what is played, but also how it is played. In this section, we evaluate content preservation (what is played) and style coherence (how it is played) using ob- jective metrics. Content preservation assesses whether the melody and harmony are well-maintained. We use Melody Chroma Accuracy (MCA) [32, 33] and Chord Accuracy (CA) [47, 48] to measure these aspects. Style coherence evaluates whether the accompaniment grooves and per- formance dynamics are well captured and manifested in the symbolic arrangement. We introduce three metrics: Grooving Pattern Coherence (GPC) [49], Velocity Con- tour Coherence (VCC), and Tempo Accuracy (TA). Among them, GPC and VCC compare the generated covers to hu- man arrangements (available for POP909). TA compares the generated tempo to the ground-truth (estimatable from audio using [44] when not annotated). Taken together, these metrics indicate how well the generated cover cap- tures the stylistic “feel” of the reference audio. Across these metrics, a higher value indicates better coherence. Detailed metric definitions are provided in Appendix B. We consider two evaluation settings: 1) in-distribution CoherenceNaturalnessCreativityMusicality 1.0 2.0 3.0 4.0 OursPCG2w/o PTA2M Figure 3: Subjective evaluation on music quality. evaluation on the POP909 test set, and 2) out-of- distribution evaluation on 100 tracks randomly drawn from the Ballroom and GTZAN datasets. We run our method and baseline models at each test piece in 10 indepen- dent rounds, deriving 450 sets of piano cover samples for POP909, and 1000 sets for Ballroom/GTZAN. We report the mean and standard errors. As shown in Table 1, while our method leads in most metrics, we find that on POP909, it achieves a lower content preservation score (MCA and CA) than the PCG2 baseline. This is probably because our arrangement process relies on Sheetsage-extracted lead sheets, where wrongly estimated chords or melody notes at this stage can propagate to our final results. However, on the out-of-distribution Ballroom/GTZAN dataset, our model surpasses PCG2 in MCA and CA, suggesting that PCG2 is more tailored to pop music (where chords and melodies are relatively simple), whereas our approach gen- eralizes better across diverse music genres and instrumen- tations. In terms of style coherence, our model outperforms all baselines in GPC and VCC on POP909, and in TA on both test sets, indicating that it can better capture the stylis- tic characteristics from the audio. 4.4 Subjective Evaluation We further conduct a double-blind online listening sur- vey to evaluate the music quality. The survey comprises 6 test pieces of varied genres drawn from the Ballroom and the GTZAN datasets. Each test piece is accompa- nied by 4 piano covers interpreted by our model and each baseline model. For each model, we select the best result from 3 generated samples. The purpose is to prevent occa- sional low-probability failures (e.g., incomplete or degen- erate generations) from disproportionately influencing the subjective assessment. All samples are 16 bars long and rendered to audio using the Cakewalk TTS-1 soundfont, resulting in approximately 40s of audio per sample. Both the order of the test pieces and the order of samples are ran- domized. Participants are asked to complete 3 test pieces by rating each piano cover on a 5-point Likert scale across 4 criteria: 1) Audio-to-Symbolic Coherence, 2) Natural- ness, 3) Creativity, and 4) Overall Musicality. More details of the subjective evaluation are provided in Appendix C. A total of 21 participants with diverse musical back- grounds completed our survey. The average completion time is 12 minutes. Figure 3 shows the mean ratings and standard errors analyzed using repeated-measures (within- Table 2: Objective evaluation on audio-to-symbolic style transfer. The star (*) indicates significant differences (p < 0.05) based on the Friedman test. MCACAGPCVCCTA Ours28.9±0.5*40.0±1.0*64.0±1.162.8±1.173.8±2.0* w/o PT25.1±0.533.2±1.062.9±1.362.0±1.369.2±2.1 subject) ANOVA [50]. The results reveal significant main effects (p-value p < 0.05) across all evaluation criteria. While our model performs comparably to the state-of-the- art PCG2 in Naturalness, it rates higher on the remaining criteria. A Bonferroni post-hoc test further confirms that our model significantly outperforms all baselines in Coher- ence and Musicality. These results align with the objective evaluation and demonstrate that our model captures music style more effectively and produces coherent, high-quality piano cover arrangements. 5. ADDITIONAL EVALUATIONS In this section, we explore additional experimental set- tings to further evaluate our model’s capabilities, with a particular focus on cross-modal representation learning. Specifically, we examine cross-modal style transfer in Sec- tion 5.1, and audio-to-symbolic retrieval in Section 5.2. 5.1 Evaluation on Cross-Modal Style Transfer To the best of our knowledge, our method is the first to en- able cross-modal audio-to-symbolic style transfer, and thus no established baseline exists for direct comparison. We provide style-varied arrangement demo in Appendix D. In addition, we conduct an objective ablation study compar- ing our full model (Ours) against a variant without Stage-I pre-training (w/o PT ), to demonstrate the validity of our two-stage training paradigm for cross-modal style transfer. For this experiment, we collect content lead sheets from the POP909 test split. In 10 independent trials, each lead sheet is paired with a style audio reference randomly drawn from the PIAST test split. This yields 450 style-transfer pairs, and both our full model and the ablation variant gen- erate outputs for each pair. We evaluate content preser- vation using MCA and CA, which measure melody/chord similarity to the input lead sheet, and style coherence using GPC, VCC, and TA, which measure groove/velocity/tempo alignment with the style reference audio’s transcription. In Table 2, our full model wins the w/o PT variant across all metrics, significantly so for MCA, CA, and TA, indicating stronger style-transfer capability in both content preservation and style coherence, thereby validating our two-stage training paradigm. This finding also aligns with the subjective evaluation results in Section 4.4, where our full model is consistently preferred in terms of Audio-to- Symbolic Coherence, besides other musicality criteria. 5.2 Evaluation on Audio-to-Symbolic Alignment While the Q-Former bridges the modality gap for audio-to- symbolic arrangement, it can also operate independently Table 3: Objective evaluation on audio-to-MIDI retrieval. Results are compared against baseline models and across two evaluation settings: MIDI with random transposition (right) and without transposition (left), highlighting robustness in capturing stylistic features beyond absolute pitch and key. w/o Transpositionw/ Random Transposition Acc@1 (%)↑Acc@5 (%)↑Rank↓Acc@1 (%)↑Acc@5 (%)↑Rank↓ Random1.4± 0.34.8± 0.864.4± 1.70.7± 0.24.2± 0.565.4± 1.0 CLaMP3.4± 0.415.0± 0.642.5± 0.23.4± 0.411.0± 0.648.5± 0.3 Ours71.4± 1.595.1± 0.52.1± 0.170.2± 1.594.8± 0.52.1± 0.1 as an audio-to-symbolic retriever. In this setting, the Q- Former measures stylistic coherence between an audio clip and a symbolic segment by comparing their learned repre- sentations. Given an audio query and a set of symbolic candidates, the model can retrieve the symbolic piece that best aligns with the music style of the query audio. In this section, we compare our approach with CLaMP3 [51] to assess whether our method achieves superior performance. 5.2.1 Audio-to-MIDI Retrieval To assess the alignment capability of the Q-Former, we evaluate it after Stage-I training on the audio-to-MIDI re- trieval task. In each of 10 independent runs, we construct a test set of 128 pairs of 10s audio and 4-bar MIDI, ran- domly sampled from PIAST and POP909 (64 pairs each). For each audio query, the model is tasked with retrieving its corresponding MIDI from the full pool of 128 candi- dates. We consider two evaluation settings: one in which MIDI candidates are randomly transposed to all 12 keys, and the other without transposition. This setup helps us ex- amine the model’s robustness in capturing stylistic features beyond absolute pitch and key. Performance is measured using three metrics: Top-1 Accuracy (Acc@1), Top-5 Ac- curacy (Acc@5), and Mean Rank. We report the mean and standard error across the 10 resampled runs. As shown in Table 3, we compare our model against CLaMP3 in addition to a random guessing bot for san- ity check. In CLaMP3, audio and MIDI are aligned indi- rectly via text due to the greater availability of music–text pairs in both domains. While this indirect alignment al- lows CLaMP3 to perform substantially better than random guessing, its Acc@1 remains low. Interestingly, transpos- ing the MIDI candidates has a noticeable effect, leading to a 4-point drop in Acc@5 and a 6-rank increase in Mean Rank. This is probably because key signatures are fre- quently referenced in text descriptions of music, making CLaMP3 particularly sensitive to pitch-level features while struggling to capture finer stylistic nuances. In compari- son, our model consistently outperforms CLaMP3 across all metrics and exhibits negligible performance differences between the transposed and non-transposed settings. This demonstrates that our Q-Former learns more robust audio- to-symbolic alignments, effectively capturing stylistic co- herence beyond surface-level attributes. 5.2.2 Ablation Study on Pre-Training Objectives We are also interested in the contribution of each pre- training objective in Section 3.3 to cross-modal align- Table 4: Ablation study on individual pre-training objec- tives for cross-modal alignment. PIASTPOP909 Acc@1↑Rank↓Acc@1↑Rank↓ C96.7± 0.31.8± 0.236.2± 1.35.2± 0.3 C+M97.2± 0.41.6± 0.240.8± 0.95.0± 0.4 C+M+G97.1± 0.31.7± 0.244.7± 1.44.7± 0.4 ment. We conduct an ablation study on the Q-Former’s audio-to-MIDI retrieval performance based on three differ- ent pre-training configurations: contrastive loss only (C), contrastive + matching losses (C+M), and contrastive + matching + generative losses (C+M+G). As in the previ- ous section, we repeat our experiment over 10 independent runs on 128 resampled audio-MIDI pairs. As shown in Table 4, we conduct evaluation separately on PIAST and POP909. The former involves piano-only music, while the latter includes multi-instrumental accom- paniments, requiring the model to extract style from richer audio textures. We observe that the performance difference is relatively small on PIAST, suggesting that contrastive learning alone may suffice for simpler piano alignment. However, on POP909, we see both the matching and gen- erative losses contribute meaningfully to an improved re- trieval accuracy and a lower mean rank. These findings indicate that all three objectives are important for learning robust, generalizable cross-modal alignment. 6. CONCLUSION In this paper, we introduce a cross-modal framework for audio-to-symbolic arrangement. By re-purposing the Q- Former to align audio and symbolic modalities, our model extracts and applies implicit music style using pre-trained music LMs, enabling expressive piano arrangement con- ditioned on both a lead sheet and an audio reference. Through a two-stage training process—combining repre- sentation learning and generative modeling—we extract stylistic features from a frozen, large audio LM and guide a symbolic LM without re-training either backbone. We conduct quantitative experiments on piano cover genera- tion and provide qualitative demos of style transfer. Re- sults demonstrate improved audio-to-symbolic coherence and musicality, highlighting the potential of this frame- work for controllable, style-aware music generation be- yond explicitly labeled content. 7. REFERENCES [1] R. Yang, D. Wang, Z. Wang, T. Chen, J. Jiang, and G. Xia, “Deep music analogy via latent representation disentanglement,” in Proceedings of the 20th Interna- tional Society for Music Information Retrieval Confer- ence, ISMIR 2019, 2019, p. 596–603. [2] Z. Wang, D. Wang, Y. Zhang, and G. Xia, “Learning in- terpretable representation for controllable polyphonic music generation,” in Proceedings of the 21st Interna- tional Society for Music Information Retrieval Confer- ence, ISMIR 2020, 2020, p. 662–669. [3] P. Lu, X. Xu, C. Kang, B. Yu, C. Xing, X. Tan, and J. Bian, “Musecoco: Generating symbolic music from text,” arXiv preprint arXiv:2306.00110, 2023. [4] K. Bhandari, A. Roy, K. Wang, G. Puri, S. Colton, and D. Herremans, “Text2midi: Generating symbolic mu- sic from captions,” in AAAI-25, Sponsored by the Asso- ciation for the Advancement of Artificial Intelligence. AAAI Press, 2025, p. 23 478–23 486. [5] M. Wei, M. Freeman, C. Donahue, and C. Sun, “Do music generation models encode music theory?” in Proceedings of the 25th International Society for Mu- sic Information Retrieval Conference, ISMIR 2024, 2024, p. 680–687. [6] W. Ma and G. Xia, “Exploring the internal mechanisms of music llms: A study of root and quality via probing and intervention techniques,” in ICML 2024 Workshop on Mechanistic Interpretability, 2024. [7] W. Ma, X. Li, and G. Xia, “Do music llms learn sym- bolic concepts? a pilot study using probing and inter- vention,” in Audio Imagination: NeurIPS 2024 Work- shop AI-Driven Speech, Music, and Sound Generation, 2024. [8] M. A. V. Vásquez, C. Pouw, J. A. Burgoyne, and W. H. Zuidema, “Exploring the inner mechanisms of large generative music models,” in Proceedings of the 25th International Society for Music Information Retrieval Conference, ISMIR 2024, 2024, p. 791–798. [9] R. Castellon, C. Donahue, and P. Liang, “Codified au- dio language modeling learns useful representations for music information retrieval,” in Proceedings of the 22nd International Society for Music Information Re- trieval Conference, ISMIR 2021, 2021, p. 88–96. [10] L. Lin, G. Xia, J. Jiang, and Y. Zhang, “Content-based controls for music large language modeling,” in Pro- ceedings of the 25th International Society for Music Information Retrieval Conference, ISMIR 2024, 2024, p. 783–790. [11] Y. Zhang, Y. Ikemiya, W. Choi, N. Murata, M. A. Martínez-Ramírez, L. Lin, G. Xia, W.-H. Liao, Y. Mit- sufuji, and S. Dixon, “Instruct-musicgen: Unlocking text-to-music editing for music language models via instruction tuning,” arXiv preprint arXiv:2405.18386, 2024. [12] L. Lin, G. Xia, Y. Zhang, and J. Jiang, “Arrange, in- paint, and refine: Steerable long-term music audio gen- eration and editing via content-based controls,” in Pro- ceedings of the Thirty-Third International Joint Con- ference on Artificial Intelligence, IJCAI 2024.ij- cai.org, 2024, p. 7690–7698. [13] S.-L. Wu, C. Donahue, S. Watanabe, and N. J. Bryan, “Music controlnet: Multiple time-varying controls for music generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, p. 2692– 2703, 2024. [14] C. Donahue, J. Thickstun, and P. Liang, “Melody tran- scription via generative pre-training,” in Proceedings of the 23rd International Society for Music Information Retrieval Conference, ISMIR 2022, 2022, p. 485–492. [15] Z. Wang, D. Xu, G. Xia, and Y. Shan, “Audio-to- symbolic arrangement via cross-modal music repre- sentation learning,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022. IEEE, 2022, p. 181–185. [16] J. Li, D. Li, S. Savarese, and S. C. H. Hoi, “BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,” in In- ternational Conference on Machine Learning, ICML 2023, ser. Proceedings of Machine Learning Research, vol. 202. PMLR, 2023, p. 19 730–19 742. [17] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar et al., “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023. [18] J. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Shar- ifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zis- serman, and K. Simonyan, “Flamingo: a visual lan- guage model for few-shot learning,” in Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, 2022. [19] Z. Kong, A. Goel, R. Badlani, W. Ping, R. Valle, and B. Catanzaro, “Audio flamingo: A novel audio lan- guage model with few-shot learning and dialogue abil- ities,” in Forty-first International Conference on Ma- chine Learning, ICML 2024, Vienna, Austria, July 21- 27, 2024. OpenReview.net, 2024. [20] R. Yuan, H. Lin, S. Guo, G. Zhang, J. Pan, Y. Zang, H. Liu, Y. Liang, W. Ma, X. Du et al., “Yue: Scaling open foundation models for long-form music genera- tion,” arXiv preprint arXiv:2503.08638, 2025. [21] A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi et al., “Musiclm: Generatingmusicfromtext,”arXivpreprint arXiv:2301.11325, 2023. [22] J. Melechovský, Z. Guo, D. Ghosal, N. Majumder, D. Herremans, and S. Poria, “Mustango: Toward con- trollable text-to-music generation,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Hu- man Language Technologies (Volume 1: Long Papers), NAACL 2024.Association for Computational Lin- guistics, 2024, p. 8293–8316. [23] J. Thickstun, D. L. W. Hall, C. Donahue, and P. Liang, “Anticipatory music transformer,” Transactions on Machine Learning Research, 2024. [24] J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Syn- naeve, Y. Adi, and A. Défossez, “Simple and control- lable music generation,” in Advances in Neural Infor- mation Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, 2023. [25] Q. Kong, B. Li, X. Song, Y. Wan, and Y. Wang, “High- resolution piano transcription with pedals by regress- ing onset and offset times,” IEEE ACM Trans. Audio Speech Lang. Process., vol. 29, p. 3707–3717, 2021. [26] L. Ou, Z. Guo, E. Benetos, J. Han, and Y. Wang, “Exploring transformer’s potential on automatic pi- ano transcription,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022. IEEE, 2022, p. 776–780. [27] J. Gardner, I. Simon, E. Manilow, C. Hawthorne, and J. H. Engel, “MT3: multi-task multitrack music tran- scription,” in The Tenth International Conference on Learning Representations, ICLR 2022.OpenRe- view.net, 2022. [28] X. Gu, L. Ou, W. Zeng, J. Zhang, N. Wong, and Y. Wang, “Automatic lyric transcription and automatic music transcription from multimodal singing,” ACM Trans. Multim. Comput. Commun. Appl., vol. 20, no. 7, p. 209:1–209:29, 2024. [29] W. Zeng, X. He, and Y. Wang, “End-to-end real-world polyphonic piano audio-to-score transcription with hi- erarchical decoding,” in Proceedings of the Thirty- Third International Joint Conference on Artificial In- telligence, IJCAI 2024.ijcai.org, 2024, p. 7788– 7795. [30] E. Nakamura and K. Yoshii, “Statistical piano re- duction controlling performance difficulty,” APSIPA Transactions on Signal and Information Processing, vol. 7, 2018. [31] C. Tan, S. Guan, and Y. Yang, “Picogen: Generate pi- ano covers with a two-stage approach,” in Proceedings of the 2024 International Conference on Multimedia Retrieval, ICMR 2024. ACM, 2024, p. 1180–1184. [32] C. Tan, H. Ai, Y. Chang, S. Guan, and Y. Yang, “Pico- gen2: Piano cover generation with transfer learning ap- proach and weakly aligned data,” in Proceedings of the 25th International Society for Music Information Re- trieval Conference, ISMIR 2024, 2024, p. 555–562. [33] J. Choi and K. Lee, “Pop2piano : Pop audio-based piano cover generation,” in IEEE International Con- ference on Acoustics, Speech and Signal Processing ICASSP 2023. IEEE, 2023, p. 1–5. [34] J. Li, R. R. Selvaraju, A. Gotmare, S. R. Joty, C. Xiong, and S. C. Hoi, “Align before fuse: Vision and lan- guage representation learning with momentum distil- lation,” in Advances in Neural Information Processing Systems 34: Annual Conference on Neural Informa- tion Processing Systems 2021, NeurIPS 2021, 2021, p. 9694–9705. [35] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in The Tenth In- ternational Conference on Learning Representations, ICLR 2022. OpenReview.net, 2022. [36] Z. Wang, K. Chen, J. Jiang, Y. Zhang, M. Xu, S. Dai, and G. Xia, “POP909: A pop-song dataset for music arrangement generation,” in Proceedings of the 21st International Society for Music Information Retrieval Conference, ISMIR 2020, 2020, p. 38–45. [37] H. Bang, E. Choi, M. Finch, S. Doh, S. Lee, G.-H. Lee, and J. Nam, “Piast: A multimodal piano dataset with audio, symbolic and text,” in Proceedings of the 3rd Workshop on NLP for Music and Audio (NLP4MusA), 2024, p. 5–10. [38] F. Gouyon, A. Klapuri, S. Dixon, M. Alonso, G. Tzane- takis, C. Uhle, and P. Cano, “An experimental compari- son of audio tempo induction algorithms,” IEEE Trans. Speech Audio Process., vol. 14, no. 5, p. 1832–1844, 2006. [39] F. Krebs, S. Böck, and G. Widmer, “Rhythmic pattern modeling for beat and downbeat tracking in musical audio,” in Proceedings of the 14th International Soci- ety for Music Information Retrieval Conference, ISMIR 2013, 2013, p. 227–232. [40] G. Tzanetakis and P. R. Cook, “Musical genre classi- fication of audio signals,” IEEE Trans. Speech Audio Process., vol. 10, no. 5, p. 293–302, 2002. [41] U. Marchand and G. Peeters, “Swing ratio estimation,” in Proceedings of the 18th International Conference on Digital Audio Effects, DAFx-15, 2015, p. 1–6. [42] C. Donahue and P. Liang, “Sheet sage: Lead sheets from music audio,” ISMIR 2021 Late-Breaking and Demo, 2021. [43] P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever, “Jukebox: A generative model for music,” arXiv preprint arXiv:2005.00341, 2020. [44] S. Böck and M. E. P. Davies, “Deconstruct, analyse, reconstruct: How to improve tempo, beat, and down- beat estimation,” in Proceedings of the 21th Interna- tional Society for Music Information Retrieval Confer- ence, ISMIR 2020, 2020, p. 574–582. [45] J. Jiang, K. Chen, W. Li, and G. Xia, “Large- vocabulary chord transcription via chord structure de- composition,” in Proceedings of the 20th International Society for Music Information Retrieval Conference, 2019, p. 644–651. [46] C. Hawthorne, E. Elsen, J. Song, A. Roberts, I. Si- mon, C. Raffel, J. H. Engel, S. Oore, and D. Eck, “On- sets and frames: Dual-objective piano transcription,” in Proceedings of the 19th International Society for Music Information Retrieval Conference, ISMIR 2018, 2018, p. 50–57. [47] Y. Ren, J. He, X. Tan, T. Qin, Z. Zhao, and T.-Y. Liu, “Popmag: Pop music accompaniment generation,” in Proceedings of the 28th ACM International Conference on Multimedia, 2020, p. 1198–1206. [48] J. Zhao, G. Xia, Z. Wang, and Y. Wang, “Structured multi-track accompaniment arrangement via style prior modelling,” in Advances in Neural Information Pro- cessing Systems 38: Annual Conference on Neural In- formation Processing Systems 2024, NeurIPS 2024, 2024. [49] S. Wu and Y. Yang, “The jazz transformer on the front line: Exploring the shortcomings of ai-composed mu- sic through quantitative measures,” in Proceedings of the 21st International Society for Music Information Retrieval Conference, 2020, p. 142–149. [50] H. Scheffe, The analysis of variance.John Wiley & Sons, 1999, vol. 72. [51] S. Wu, Z. Guo, R. Yuan, J. Jiang, S. Doh, G. Xia, J. Nam, X. Li, F. Yu, and M. Sun, “Clamp 3: Universal music information retrieval across unaligned modali- ties and unseen languages,” in Findings of the Associa- tion for Computational Linguistics, ACL 2025. Asso- ciation for Computational Linguistics, 2025, p. 2605– 2625. [52] M. Zeng, X. Tan, R. Wang, Z. Ju, T. Qin, and T. Liu, “Musicbert: Symbolic music understanding with large- scale pre-training,” in Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, ser. Findings of ACL, vol. ACL/IJCNLP 2021.Associa- tion for Computational Linguistics, 2021, p. 791–800. [53] Y.-S. Huang and Y.-H. Yang, “Pop music transformer: Beat-based modeling and generation of expressive pop piano compositions,” in M ’20: The 28th ACM Inter- national Conference on Multimedia, 2020, p. 1180– 1188. [54] I. Loshchilov and F. Hutter, “Decoupled weight de- cay regularization,” in 7th International Conference on Learning Representations, ICLR 2019.OpenRe- view.net, 2019. [55] C. Raffel, B. McFee, E. J. Humphrey, J. Salamon, O. Nieto, D. Liang, and D. P. W. Ellis, “Mir_eval: A transparent implementation of common MIR metrics,” in Proceedings of the 15th International Society for Music Information Retrieval Conference, ISMIR 2014, 2014, p. 367–372. [56] S. Rouard, F. Massa, and A. Défossez, “Hybrid trans- formers for music source separation,” in IEEE Inter- national Conference on Acoustics, Speech and Signal Processing ICASSP 2023. IEEE, 2023, p. 1–5. [57] M. Mauch and S. Dixon, “PYIN: A fundamental fre- quency estimator using probabilistic threshold distribu- tions,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2014.IEEE, 2014, p. 659–663. [58] B. McFee, C. Raffel, D. Liang, D. P. W. Ellis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and music signal analysis in python,” in Pro- ceedings of the 14th Python in Science Conference 2015 (SciPy 2015). scipy.org, 2015, p. 18–24. [59] A. L. Uitdenbogerd and J. Zobel, “Manipulation of music for melody matching,” in Proceedings of the 6th ACM International Conference on Multimedia ’98. ACM, 1998, p. 235–240. [60] J. Jiang,“MIDI Chord Recognition via Bar- LevelModeling,”https://github.com/music-x-lab/ midi-chord-recognition, 2025. [61] S. Dixon, “Automatic extraction of tempo and beat from expressive performances,” Journal of New Music Research, vol. 30, no. 1, p. 39–58, 2001. A. MODEL AND TRAINING DETAILS Our model comprises three components: an audio LM, a symbolic music LM, and a Q-Former connecting the two. This section provides detailed configurations of each com- ponent module and the training details. A.1 Q-Former The Q-Former is initialized with the MusicBERT-Base model [52]. The added cross-attention layers are randomly initialized. Following BLIP-2 [16], we use K = 32 learn- able queries, each with dimension 768. Symbolic piano arrangements are tokenized in the OctMIDI format [52], which produces note-wise joint embeddings.Overall, the Q-Former comprises 186M parameters, including the learnable queries and note embedding layers. A.2 Audio LM We use MusicGen-Large [24] as our audio LM. We discard the text encoder and retain only the music decoder, a 48- layer Transformer. Audio codecs are fed to the decoder and we extract the hidden representations from the 25th layer, as prior probing studies [5, 7–9] suggest that middle layers capture more musically meaningful features. This setup retains 1.7B frozen parameters from MusicGen. A.3 Symbolic Music LM For symbolic arrangement,we adopt MuseCoco- xLarge [3], which is a 24-layer Transformer decoder pre-trained on large-scale symbolic music corpora. We remove its text-related components and keep 1.2B frozen parameters from the music decoder. Symbolic note tokens are converted into the REMI format [53]. Despite the slightly different tokenizations used across stages, we find that the latent representations learned in Stage-I remain compatible with Stage-I. To enable compatibility with MuseCoco, we projectZ into the same embedding dimension as MuseCoco’s token embeddings via a linear layer. A LoRA adapter with rank 16 accommodates the added lead sheet condition. A.4 Training Details Model training focuses on the 186M Q-Former parame- ters, which is significantly smaller than the billion-scale LM backbones. In Stage-I, the Q-Former is pre-trained in FP16 using batch size 128 for 10 epochs (130K itera- tions). The LoRA adapter in Stage-I adds 5M parame- ters and we fine-tune the model for another 5 epochs using batch size 32. Both training stages are conducted on four NVIDIA A40 GPUs (48GB each). We use the AdamW optimizer [54] with an initial learning rate of 1e-4, a lin- ear warm-up over the first 1K steps, and a cosine decay schedule to a final rate of 1e-5. At test time, we use top-k sampling with k = 15. B. OBJECTIVE METRICS We introduce five statistical metrics to evaluate content preservation and style coherence for the piano cover gen- eration tasks. This section provides the definitions. B.1 Melody Chroma Accuracy (MCA) We use the melody.raw_chroma_accuracy metric 2 pro- vided by mir_eval [55] to evaluate the similarity between two monophonic melody sequences. For the reference melody, we apply Demucs [56] to isolate the vocal stem from the audio and then extract the F0 contour using pYIN [57] provided by librosa [58]. For the estimated melody, we obtain the melody skyline [59] from the gen- erated piano cover MIDI and convert the MIDI pitches to frequencies. The two melodies are compared position-wise under a tolerance of 50 cents. While this implementation follows [32], we additionally ensure that the two melodies are temporally aligned to the same sequence length. B.2 Chord Accuracy (CA) We introduce Chord Accuracy from [47, 48] to measure the similarity between two chord sequences. For the refer- ence sequence, when annotated chords are not available, we apply the method of [45] to detect chords from the input audio. For the estimated sequence, we use [60] to detect chords from the generated piano cover MIDI. Both chord sequences are aligned and compared at 1-beat granu- larity in terms of root and full quality based on the MIREX tetrads rule. 3 B.3 Grooving Pattern Coherence (GPC) Grooving Pattern Coherence (GPC) evaluates the groov- ing pattern similarity between the generated piano cover and human’s arrangement. The grooving pattern, which is defined in [49], represents the positions in a MIDI segment at which there is at least one note onset. We consider 4-bar segments at 1/4-beat granularity, deriving grooving pattern g as a 64-dimensional binary vector. The GPC over a test piece is defined as follows: GPC = 1 N N X n=1 cos(g hm n ,g cv n ),(1) where N is the number of non-overlapping 4-bar segments in the test piece. cos(·,·) computes the cosine similarity. g hm andg cv represent the grooving pattern feature from the human arrangement and the generated piano cover, re- spectively. The GPC metric is omitted for evaluation on Ballroom/GTZAN, where human’s piano arrangement is not available. 2 https://mir-eval.readthedocs.io/latest/api/melody.html #mir_eval.melody.raw_chroma_accuracy 3 https://mir-eval.readthedocs.io/latest/api/chord.html #mir_eval.chord.tetrads B.4 Velocity Contour Coherence (VCC) Velocity Contour Coherence (VCC) evaluates the similar- ity of the velocity contour between the generated piano cover and the human arrangement, using the same formu- lation as GPC. We define the velocity contour as a time- series feature representing the average note velocity at each timestep. It has the same dimensionality and temporal granularity as the grooving pattern, but instead of a binary vector, it consists of real values ranging from 0 to 127. The VCC metric is omitted for Ballroom/GTZAN, where human’s piano arrangement is not available. B.5 Tempo Accuracy (TA) Tempo Accuracy (TA) evaluates the correctness of the esti- mated tempo tp est relative to the reference (ground-truth) tempo tp ref . We first define correctness indicator D(k) for one test piece as follows: D(k) = 1 |tp ref − k× tp est | tp ref < 0.08 ,(2) where 1(· ) is the indicator function. k ∈ 1 ⁄ 2 , 1, 2 takes into account the octave ambiguity.The toler- ance threshold of 0.08 follows the empirical setting in mir_eval. 4 We then define tempo accuracy TA as: TA = max D(1), 0.5· D( 1 ⁄ 2 ), 0.5· D(2) ,(3) where we consider half-tempo and double-tempo matches as partially correct (weight 0.5). This is because tempo perception is known to exhibit octave ambiguity, where perceptions at multiple metrical levels are still considered valid rather than true perceptual errors [61]. We derive the ground-truth tempo from beat annotations when available. On Ballroom/GTZAN, we use the audio tempo estimation results by madmom [44], and we omit the A2M baseline in this setting because it already relies on madmom’s estimation in its pipeline. C. SUBJECTIVE EVALUATION DETAILS Our subjective evaluation is conducted through an online crowdsourcing study, where participants complete a sur- vey consisting of listening and rating tasks. This section provides additional details on the survey design. C.1 General Instructions Participants receive the general instructions, which clarify their rights and the conditions of participation: • Participation is entirely voluntary, and one may withdraw at any time without any negative conse- quences. • No personally identifying information is collected; all responses are anonymous and used solely for re- search purposes. 4 https://mir-eval.readthedocs.io/latest/api/tempo.html #mir_eval.tempo.detection C.2 Survey Design Our survey consists of 6 pages, each presenting 4 versions of piano cover arrangements corresponding to a common test audio piece. The test audio pieces are drawn from the Ballroom and GTZAN datasets and span a variety of genres. The 4 arrangement versions are produced by our model and all baseline models (PCG2, A2M, and w/o PT ), respectively. For each model, we select the best result from 3 independently generated samples to avoid occa- sional low-probability failures (e.g., incomplete or degen- erate outputs) from disproportionately affecting subjective evaluation. All models to be evaluated are anonymized, and their orders on each page are randomized. We acknowledge that participants may use different in- ternal criteria when evaluating music. To minimize such variability, we conduct within-subject ANOVA [50] for data analysis, which ensures that variances in ratings re- flect differences among the models rather than participants. To control response quality, we only accept complete eval- uation sets; that is, participants have to rate all four models under a common test piece for their responses to be con- sidered valid. C.3 Completion Time Each participant is randomly assigned 3 out of the 6 pages. On each page, participants first listen to the test audio piece, and then listen to and rate 4 corresponding piano cover samples. All samples in the survey are 16 bars long and rendered to audio using the Cakewalk TTS-1 sound- font, producing approximately 40 seconds of audio per sample. This design targets a total completion time of 10–15 minutes, ensuring that participants have sufficient time to listen carefully without excessive fatigue. The ac- tual completion time observed on average is 12 minutes. C.4 Participant Profiles Participants are asked to self-identify their musical back- ground as amateur, intermediate, or professional, follow- ing the guidelines below: • Amateur:I enjoy listening to music.I can play/sing/compose short music pieces. I know a lit- tle music theory. I can evaluate a composition based on my feelings. • Intermediate: I have some experience in perform- ing, composing, or other music activities. I know a certain amount of music theories that can help me evaluate a composition. • Professional: I am now pursuing/have completed a music degree, or having equivalent background. I am proficient in using music theory to evaluate a composition. Among the 21 valid responses collected, 6 participants identified as amateur (28.6%), 11 as intermediate (52.4%), and 4 as professional (19.0%). All authors are excluded from the survey. FB♭B♭dimD♭m7♭5FDm7FB♭dimG7B♭B♭dimB♭C7 (a) The input lead sheet. = 122 (b) Piano cover based on the original soundtrack. = 81 (c) Piano arrangement in ragtime. = 114 (d) Piano arrangement in bossa nova. Figure 4: Audio-to-symbolic arrangement for an 8-bar excerpt from The Sound of Music. Score is manually engraved from MIDI for visualization purpose. Figures 4b to 4d are arranged based on the lead sheet in 4a and an audio reference from the original soundtrack, a ragtime piece, and a bossa nova piece, respectively. Preserved music contents are highlighted in blue note heads. Synthesized audio is provided on the demo page: https://zhaojw1998.github.io/bossa/. D. ARRANGEMENT DEMONSTRATION In this section, we demonstrate the performance of our audio-to-symbolic arrangement model under freely manip- ulated audio style references. Figure 4a shows an 8-bar lead sheet excerpt from the musical The Sound of Music. The selected passage features harmonically rich chords, including diminished and seventh chord qualities, which present suitable complexity for arrangement experiments. Figures 4b to 4d showcase the arrangement results condi- tioned on varied audio references. The 8-bar arrangement is generated using windowed sampling, wherein a 4-bar context window progresses forward every 2 bars and con- tinues sampling conditioned upon the preceding 2 bars. Figure 4b shows the piano cover from the original The Sound of Music soundtrack, 5 which features lush or- chestration dominated by string ensembles. Our arrange- ment captures this orchestral essence through dense, block- chord voicing that emulates the sonority of string sec- tions. Additionally, ornaments such as arpeggios and trills are found to complement the sweeping harmonic textures, which contributes to the music’s free-flowing character. Figure 4c shows an arrangement conditioned on the ragtime classic The Entertainer. 6 Following the audio 5 Original audio: https://youtu.be/6f0T6UV-HiI&t=57 6 Ragtime audio: https://youtu.be/jKlfNfRZL9I&t=11 recording, the arrangement’s tempo is “not fast,” and the piano texture distinctly adopts a ragtime rhythm, with steady bass notes on downbeats and syncopated chordal accents on upbeats. Figure 4d shows an arrangement conditioned on the bossa nova piece The Girl from Ipanema. 7 In this inter- pretation, the arrangement is characterized by a moderate tempo and distinctive left-hand syncopated patterns char- acteristic of the bossa nova genre. Across all three piano arrangements, while distinct mu- sic styles are effectively captured from the audio refer- ences, the theme melody and harmonic structures remain faithfully preserved. In Figure 4, we highlight melody notes preserved from the lead sheet using blue note heads. Additional examples are available on our demo page, including arrangements in a wider range of classical, jazz, and pop styles applied to well-known lead sheets, illustrat- ing the flexibility of both style and lead sheet control. E. LIMITATIONS Our proposed method demonstrates the ability to learn implicit music style from audio.At the current stage of this work, we acknowledge that the extracted style primarily represents segment-level global characteristics. 7 Bossa nova audio: https://youtu.be/DvA_wDOVD10&t=12 Here, “global” refers to a latent style profile at the level of short 4-bar segments, which is still considerably more fine-grained (and more local) than typical global attributes such as genre labels and textual style tags. Our underly- ing assumption is that music style remains locally consis- tent at the segment or bar level. We empirically found that this assumption holds across most music traditions, which guarantees our model to perform reliably under this design. While longer generation can be achieved us- ing windowed sampling, we acknowledge that this ap- proach may smooth over intended stylistic transitions at phrase boundaries, thus leading to diminished expressiv- ity at longer timescales. When considering inter-phrase and longer-term music development, we recognize that a single segment-level style representation is insufficient to capture evolving dynamics. Also, subject to the availabil- ity of audio-symbolic data, this work is dedicated to piano arrangement. The cross-modal arrangement of long-term, multi-track music may require hierarchical or temporally adaptive style modeling, which is an important direction for our future work.