Paper deep dive
Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding
Tianle Wang, Xinyi Tong, Liangke Zhao, Jishang Chen, Sirui Zhang, Haoxin Zhang, Xin Jin, Duo Xu, Xiaobing Li, Song-Chun Zhu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/29/2026, 3:13:05 AM
Summary
The paper introduces the Dissonance Spectrum (DS), a nonnegative time-frequency representation that models perceptual frequency interactions using a tolerance-based rational pitch-relation kernel and logarithmic harmonic distance. Applied to Constant-Q Transforms (CQT), DS attributes aggregate pairwise consonance/dissonance relations back to individual frequency bins. Experiments show DS improves performance in open-ended music question answering (MU-LLaMA) and music emotion recognition (Music2Emo) compared to baselines, parameter-matched Gaussian branches, and architecture-matched magnitude-CQT branches.
Entities (8)
Relation Signals (6)
Dissonance Spectrum → isintegratedinto → MU-LLaMA
confidence 95% · add a lightweight DS encoder to MU-LLaMA for open-ended music question answering
Dissonance Spectrum → isintegratedinto → Music2Emo
confidence 95% · and to Music2Emo for categorical and dimensional emotion recognition
Dissonance Spectrum → uses → Constant-Q Transform
confidence 95% · applies a tolerance-based rational pitch-relation kernel ... to a constant-Q spectrum
Dissonance Spectrum → useskernel → Harmonic Distance
confidence 93% · applies a tolerance-based rational pitch-relation kernel with logarithmic harmonic distance
Dissonance Spectrum → improvesperformanceon → Music Question Answering
confidence 92% · DS obtains the highest mean on every reported endpoint ... in open-ended music question answering
Dissonance Spectrum → improvesperformanceon → Music Emotion Recognition
confidence 92% · DS obtains the highest mean on every reported endpoint ... in categorical and dimensional music emotion recognition
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Conventional music representations describe acoustic energy over time and frequency but do not explicitly expose relations among simultaneous frequency components. We introduce the \emph{Dissonance Spectrum} (DS), a nonnegative time--frequency representation that applies a tolerance-based rational pitch-relation kernel with logarithmic harmonic distance to a constant-Q spectrum and attributes aggregate pairwise interactions back to individual frequency bins. Controlled music-theory tests show strong ordinal agreement for intervals, harmonic-function connections, and church modes, and weaker but significant agreement across diverse chord voicings. DS is then encoded by a lightweight parallel branch whose zero-initialized residual projection preserves the baseline function at initialization. Across six paired training seeds in open-ended music question answering and categorical and dimensional music emotion recognition, DS obtains the highest mean on every reported endpoint relative to the unchanged baseline, a parameter-matched Gaussian-input branch, and an architecture-matched magnitude-CQT branch. These results support DS as an interpretable, complementary representation, while listener-specific perception and broader task coverage remain open problems.
Tags
Links
- Source: https://arxiv.org/abs/2608.25621v1
- Canonical: https://arxiv.org/abs/2608.25621v1
Trouble viewing inline? Open PDF directly →
Full Text
93,056 characters extracted from source content.
Expand or collapse full text
Beyond Frequency: Dissonance Spectrum for Perceptually Motivated Music Understanding Tianle Wang Xinyi Tong Liangke Zhao Jishang CHEN Sirui Zhang Haoxin Zhang Xin Jin Duo XU Xiaobing Li Song-Chun Zhu Abstract Conventional music representations describe acoustic energy over time and frequency but do not explicitly expose relations among simultaneous frequency components. We introduce the Dissonance Spectrum (DS), a nonnegative time–frequency representation that applies a tolerance-based rational pitch-relation kernel with logarithmic harmonic distance to a constant-Q spectrum and attributes aggregate pairwise interactions back to individual frequency bins. Controlled music-theory tests show strong ordinal agreement for intervals, harmonic-function connections, and church modes, and moderate positive agreement across diverse chord voicings. DS is then encoded by a lightweight parallel branch whose zero-initialized residual projection preserves the baseline function at initialization. Across six paired training seeds in open-ended music question answering and categorical and dimensional music emotion recognition, DS obtains the highest mean on every reported endpoint relative to the unchanged baseline, a parameter-matched Gaussian-input branch, and an architecture-matched magnitude-CQT branch. These results support DS as an interpretable, complementary representation, while listener-specific perception and broader task coverage remain open problems. 1Beijing Institute for General Artificial Intelligence, Beijing, China 2Central Conservatory of Music, Beijing, China 3Peking University, Beijing, China - Figure 1: DS converts spectral energy into a time–frequency map of modeled consonance–dissonance relations. Introduction Music understanding extends beyond detecting acoustic events: listeners also respond to relations among simultaneous frequency components. Beating, critical-band interactions, periodicity, and harmonic organization contribute to consonance, dissonance, stability, tension, and emotion (von Helmholtz 1954; Plomp and Levelt 1965; Langner 1992; Stolzenburg 2015; Harrison and Pearce 2020). Neural systems learn rich musical information from waveforms, time–frequency representations, and audio–text supervision (Lu et al. 2021; Li et al. 2024; Won et al. 2024; Huang et al. 2022; Elizalde et al. 2023; Liu et al. 2024; Deng et al. 2024). Yet spectral magnitude mainly marks active components, while learned embeddings entangle multiple attributes; neither explicitly localizes simultaneous relations under a specified consonance–dissonance model. Music-specific inductive biases can complement end-to-end learning. Chroma and tonal features expose pitch-class organization, perceptual mid-level attributes connect audio to emotion, MERT uses a CQT teacher, and consonance-aware supervision improves chord estimation (Harte et al. 2006; Korzeniowski and Widmer 2016; Aljanaki and Soleymani 2018; Chowdhury et al. 2019; Li et al. 2024; Poltronieri et al. 2025). Lightweight mode injection likewise benefits symbolic emotion recognition (Xia et al. 2026). These results motivate a relational representation that makes information already present in the signal easier to inspect and learn. We introduce the Dissonance Spectrum (Figure Beyond Frequency: Dissonance Spectrum for Perceptually Motivated Music Understanding). DS combines tolerance-based rational approximation (Stolzenburg 2015) with logarithmic harmonic distance (Tenney 1984), applies the resulting kernel across a magnitude CQT, and attributes weighted pair relations back to their time–frequency locations. It is a deterministic reorganization of the signal, not an additional observed modality, and is computed efficiently by frequency-axis correlation. We validate the operator on intervals, chord qualities, functional chord connections, and scales, then add a lightweight DS encoder to MU-LLaMA for open-ended music question answering and to Music2Emo for categorical and dimensional emotion recognition. Parameter-matched Gaussian and architecture-matched magnitude-CQT branches control for added capacity and a second pitch-resolved pathway. Six paired seeds and exploratory mechanism variants test whether gains are consistent with ordered relational structure rather than input resolution or model size alone. Our contributions are threefold: • We formulate a nonnegative time–frequency DS that localizes amplitude-weighted rational pitch relations and derive an efficient correlation implementation. • We establish controlled evidence that the kernel and audio representation reproduce predefined ordinal trends across intervals, chord qualities, functional connections, and modes. • We design lightweight plug-and-play adapters and show that DS achieves higher six-seed mean performance than the baseline, a parameter-matched Gaussian branch, and an architecture-matched CQT branch in two model families. Related Work Music representations and perceptual priors. The CQT provides a logarithmic frequency axis aligned with musical pitch, chroma folds energy into pitch classes, and tonal-space descriptors encode harmonic proximity (Brown 1991; Harte et al. 2006; Müller and Ewert 2011). Neural models learn broader musical attributes: SpecTNT separates spectral and temporal attention, MERT and MusicFM provide transferable music embeddings, and MuLan, CLAP, MU-LLaMA, and MusiLingo connect audio representations to language (Lu et al. 2021; Li et al. 2024; Won et al. 2024; Huang et al. 2022; Elizalde et al. 2023; Liu et al. 2024; Deng et al. 2024). These representations are effective for downstream tasks but do not provide a stable, directly inspectable account of local consonance or dissonance. Recent human-written music-QA evaluation further emphasizes robustness to unimodal shortcuts, motivating cautious interpretation of automatically generated references and lexical-overlap metrics (Weck et al. 2026). Music-specific priors remain complementary to learned representations. Chroma-based networks improve chord recognition; listener-rated mid-level attributes bridge audio and emotion; MERT uses a CQT teacher; and consonance-aware distances and label smoothing improve chord estimation (Korzeniowski and Widmer 2016; Aljanaki and Soleymani 2018; Chowdhury et al. 2019; Li et al. 2024; Poltronieri et al. 2025). DS follows this knowledge-guided direction but is computed continuously from audio and preserves time–frequency attribution. Consonance perception and computational models. Consonance and dissonance denote related but nonidentical acoustic, perceptual, and music-theoretical concepts (Cazden 1980; Wand 2012). Interference accounts relate sensory dissonance to beating among nearby partials and auditory critical bands (von Helmholtz 1954; Plomp and Levelt 1965; Hutchinson and Knopoff 1978). Periodicity and harmonicity accounts associate consonance with compact common periods, virtual fundamentals, or harmonic templates (Langner 1992; Parncutt 1989; Stolzenburg 2015). Many implementations reduce these relations to a scalar, which supports ranking but does not identify the frequency regions responsible for the value. Timbre can reshape consonance curves by changing partial locations and amplitudes (Sethares 2005; Marjieh et al. 2024). Comparative modeling and listener studies further indicate that roughness, periodicity, spectral structure, experience, register, and listener characteristics all contribute to consonance judgments (Cousineau et al. 2012; Harrison and Pearce 2020; Eerola and Lahdelma 2021; McDermott et al. 2016; Kaklamani and Simserides 2026). We therefore treat DS as a perceptually motivated low-level cue rather than a complete model of musical preference. Recorded music and consonance-aware music AI. Applying dissonance models to recordings is difficult because real audio contains overlapping sources, transients, noise, tuning variation, and time-varying balance. Schwär et al. (2025) estimate time-varying sensory dissonance in multitrack recordings and attribute mixture-level dissonance to tracks. Scalar roughness descriptors and learned mid-level dissonance ratings have also supported music emotion recognition (Aljanaki and Soleymani 2018; Panda et al. 2023). DS instead attributes aggregate relations to time–frequency bins, requires no stems or chord labels, and can be encoded as an additional input to pretrained music systems. Method Overview Our method has two parts. First, the Dissonance Spectrum (DS) assigns each active CQT bin its amplitude-weighted relation to the other active components. A continuous kernel combines tolerance-based rational approximation (Stolzenburg 2015) with Tenney’s harmonic distance (Tenney 1984); sampling it on the CQT grid permits efficient one-dimensional frequency correlation instead of a K×K×TK× K× T pair tensor. Second, a lightweight encoder and shape-preserving gated residual adapter inject cached DS maps into a host representation without changing the host outputs, heads, or losses. Zero initialization preserves the pretrained baseline function at the start of fine-tuning. We next define intrinsic and cross-reference DS, derive the correlation form, and describe both host integrations; full index derivations and optional temporal variants are in the supplement. Pitch-Relation Kernel Rational candidates and complexity. For a maximum numerator and denominator Q, the reduced rational candidate set is ℛQ=pq| 1≤p≤q≤Q,p,q∈ℤ+,gcd(p,q)=1.R_Q= \ pq\; |\;1≤ p≤ q≤ Q,\;p,q ^+,\; (p,q)=1 \. (1) For a reduced ratio, we use the complexity function C(pq)=log2(pq),C ( pq )= _2(pq), (2) which is the logarithmic product form of Tenney’s harmonic distance (Tenney 1984). Here Q is a finite search-resolution hyperparameter rather than a direct estimate of a listener attribute. For visualization only, simplicity is the reversed complexity scale, S(pq)=maxr∈ℛQC(r)−C(pq).S ( pq )= _r _QC(r)-C ( pq ). (3) Given two frequencies, we order their ratio as r=min(f1,f2)max(f1,f2),r= (f_1,f_2) (f_1,f_2), (4) and define the relative-error neighborhood α(r)=x∈ℝ+||x−r|≤αr.N_α(r)= \x ^+\; |\;|x-r|≤α r \. (5) We use α=.01α=.01, following the 1% relative tolerance used in Stolzenburg’s smoothed periodicity formulation (Stolzenburg 2015). Let p∗q∗=argminpq∈ℛQ|r−pq|. p^*q^*= _ pq _Q |r- pq |. (6) The pairwise value is then D(f1,f2)=minpq∈ℛQ∩α(r)C(pq),if ℛQ∩α(r)≠∅,C(p∗q∗),otherwise.D(f_1,f_2)= cases _ pq _Q _α(r)C ( pq ),&if R_Q _α(r)≠ ,\\[10.0pt] C ( p^*q^* ),&otherwise. cases (7) Thus ratios admitting a simpler approximation inside the tolerance receive a lower value; otherwise the closest candidate is used. This is a periodicity/harmonic-distance cue, not a critical-band roughness model. Continuous pitch intervals and octave folding. We map pitch to frequency by freq(pitch)=440⋅2pitch−6912freq(pitch)=440· 2 pitch-6912 (8) and evaluate Dinterval(Δp)=D(freq(pitchref+Δp),freq(pitchref)).D_interval( p)=D (freq(pitch_ref+ p),\,freq(pitch_ref) ). (9) With Dint,max=maxΔp∈[0,12)Dinterval(Δp),D_int,max= _ p∈[0,12)D_interval( p), (10) for Δp∈[0,12) p∈[0,12), the normalized function is Dnorm(Δp)=Dinterval(Δp)Dint,max∈[0,1].D_norm( p)= D_interval( p)D_int,max∈[0,1]. (11) Let δ12(Δp)=|Δp|mod12∈[0,12) _12( p)=| p| 12∈[0,12) denote the octave-folded interval magnitude. The resulting even function is D¯(Δp)=Dnorm(δ12(Δp)),D¯(0)=0. D( p)=D_norm\! ( _12( p) ), D(0)=0. (12) The continuous pitch argument allows evaluation between equal-tempered bins, while octave folding implements a pitch-chroma assumption. Because pitch height and pitch chroma can both affect perception (Wagner et al. 2022), folding is an explicit modeling choice rather than a claim that register never matters. The resulting curve is shown in the supplement. Intrinsic Dissonance Spectrum CQT representation and preprocessing. Let ∈ℝ≥0K×TX _≥ 0^K× T be a magnitude CQT and x→t=(:,t) x^\,t=X(:,t). With B bins per octave, its pitch and frequency grids are p→=[pitchmin+12kB]k=0K−1,f→=[freq(pitchk)]k=0K−1. p= [pitch_ + 12kB ]_k=0^K-1, f= [freq(pitch_k) ]_k=0^K-1. (13) Values below a fixed floor are set to zero, and a local spectral-peak mask may suppress nonpeak bins. We use excerpt-level global-maximum normalization in the downstream experiments and the context-dependent normalization specified for the controlled validation. DS and its magnitude-CQT control always share the same normalization within a comparison. Amplitude-weighted attribution. Motivated by the pairwise aggregation of spectral partials in dissonance-curve models (Sethares 1994; Sethares 2005), we use the simple amplitude-product coefficient and pair value A(x1,x2)=x1x2,D¯(Δp,x1,x2)=x1x2D¯(Δp).A(x_1,x_2)=x_1x_2, D( p;\,x_1,x_2)=x_1x_2 D( p). (14) We define the intrinsic DS element at bin k and frame t as dkt=1K∑l=0K−1xktxltD¯(pl−pk),d^t_k= 1K _l=0^K-1x^t_k\,x^t_l\, D(p_l-p_k), (15) and collect the frame vectors as =[d→ 0⋯d→T−1]∈ℝK×T.D= bmatrix d^\,0&·s& d^\,T-1 bmatrix ^K× T. (16) The factor 1/K1/K averages the K reference-bin contributions. Equation 15 differs from a frame-level scalar: every pair contributes to both participating locations through their own target factors, so the map preserves where the modeled relations occur. All terms are nonnegative; inactive bins have zero attribution. Cross-Reference Dissonance Spectrum The same operator can attribute the relation between a target spectrum x→ x and a separate reference spectrum x→(ref) x^(ref). In vector form, d→x→x→(ref)=1K(±→⋆valid,fx→(ref))⊙x→. d^\, x→ x^(ref)= 1K ( D^± _valid,f x^(ref) ) x. (17) Here ⋆valid,f _valid,f denotes the target-aligned valid frequency-axis correlation defined below. Intrinsic DS is the special case x→(ref)=x→ x^(ref)= x. A fixed tonic, chord, or spectral template can instead provide an explicit harmonic reference while the final factor x→ x preserves attribution to the target bins. Figure 2 contrasts the two constructions. Figure 2: Conceptual DS attribution. Left: intrinsic DS aggregates pairwise relations within one spectrum. Right: cross-reference DS evaluates the target spectrum against a separate reference while retaining target-bin attribution. Efficient Correlation Form Because D¯ D depends only on pitch difference, the K2K^2 pair table need not be stored. Define the length-(2K−1)(2K-1) sampled kernel m± ^±_m =D¯(12mB), = D\! ( 12mB ), (18) m m =−(K−1),…,K−1,±→∈ℝ2K−1. =-(K-1),…,K-1, D^± ^2K-1. stored in increasing m order, with center D¯(0)=0 D(0)=0. The target-aligned index convention is given in the supplement. Since D¯(pl−pk)=±→K+l−k−1, D(p_l-p_k)= D^±_K+l-k-1, (19) Equation 15 becomes d→t=1K(±→⋆valid,fx→t)⊙x→t. d^\,t= 1K ( D^± _valid,f x^\,t ) x^\,t. (20) For all frames, repeat the same kernel across time as ±=[→±→±⋯→±]∈ℝ(2K−1)×T, D^±= bmatrix D^±& D^±&·s& D^± bmatrix ^(2K-1)× T, (21) and compute =1K(±⋆valid,f)⊙.D= 1K\,( D^± _valid,fX) . (22) The correlation is one-dimensional along frequency and returns K×TK× T values. It avoids the O(K2T)O(K^2T) intermediate storage of direct pairwise attribution and can be batched with standard correlation operators. Figure 3 shows the corresponding matrix computation; the full index derivation and optional temporal variants are included in the supplement. Figure 3: Frequency-axis matrix implementation of cross-reference DS. Intrinsic DS follows by setting the reference and target CQT matrices equal. Representation Properties and Extraction Local attribution rather than a scalar score. A conventional spectrum reports energy at (k,t)(k,t); DS reports how strongly that component participates in the modeled relations of its frame. Unlike a scalar score that collapses register and can only modulate a representation globally, the K×TK× T map preserves register and instrumentation cues. A downstream encoder can therefore distinguish a narrow high-valued interaction among upper partials from a broad interaction spanning the bass and midrange. Each high value coincides with nonzero target magnitude and can be traced to the reference components selected by the shifted kernel. Nonnegativity, symmetry, and translation structure. Nonnegative magnitudes and kernel values make DS nonnegative. The ordered ratio and octave-folded difference make the pair relation symmetric, while the final target factor restores bin-specific attribution. Because coefficients depend on pitch difference rather than absolute bin index, equal intervals share coefficients across register. This translation structure permits cross-correlation and differs from an unconstrained learned two-dimensional filter: before training, the same specified relation is treated consistently throughout the CQT range, subject to octave folding. Continuous pitch grid. The observed frequency ratio is not quantized to the rational candidate set. Equation 7 selects a candidate within tolerance or the nearest candidate, and Equation 12 is sampled at 12/B12/B-semitone CQT increments. DS therefore responds to detuning, expressive variation, and inharmonic partials rather than only twelve integer pitch classes. The supplement gives qualitative microtonal and timbral examples without assigning a perceptual ranking. Offline extraction procedure. For each excerpt, we compute magnitude CQT with the shared floor, peak mask, and normalization; sample →± D^± from Q, α, and the CQT grid; and evaluate Equation 22. Finite outputs are transformed by log(1+) (1+D) and cached with the audio hash, sample rate, hop, pitch range, B, Q, α, and preprocessing flags. Training retrieves the DS segment at the baseline branch’s temporal indices, with deterministic validation and test crops. This avoids repeated spectral computation and ensures that the baseline and DS branches observe the same recording portion. Plug-and-Play Augmentation of Music Understanding Models DS supplements rather than replaces the host waveform encoder or task representation. Let (l)∈ℝN×dH^(l) ^N× d be the hidden sequence, or a pooled token when N=1N=1, at insertion layer l. A lightweight encoder maps the cached feature to temporal tokens, =Eϕ(log(1+))∈ℝM×ds,Z=E_φ\! ( (1+D) ) ^M× d_s, (23) using convolutional frequency compression followed by temporal convolution or attention. DS tokens share the baseline branch’s excerpt indices and valid-frame mask. A shape-preserving adapter lets the module attach to host representations of different dimensions. The baseline tokens query the DS tokens and receive a gated residual update: =softmax(((l)Q)(K)⊤da)(V), =softmax\! ( (H^(l)W_Q)(ZW_K) d_a )(ZW_V), (24) =σ(G[(l);]+G), =σ\! (W_G[H^(l);C]+b_G ), (25) ~(l) H^(l) =(l)+⊙(O). =H^(l)+G (CW_O). (26) When N=1N=1, the same equations give a pooled adapter. The unchanged output shape preserves the decoder, heads, losses, and evaluation. We initialize OW_O to zero and Gb_G negatively, so ~(l)=(l) H^(l)=H^(l) initially. Disabling the branch recovers the baseline checkpoint without conversion. Independent extraction, shape-preserving insertion, and exact fallback define the plug-and-play property. For MU-LLaMA, the 576-bin map is pooled to 144 channels and processed by three convolutional blocks, a 128-dimensional temporal encoder, depthwise temporal convolution, and single-head attention. The pooled adapter follows the original audio projection while MERT and LLaMA remain frozen. For Music2Emo, temporal DS tokens are queried after the unchanged 512-dimensional projection of MERT, chord, and key-mode features, before the original emotion heads. Figure 4 shows both insertion paths and frozen versus trainable modules. Figure 4: Plug-and-play DS integration in Music2Emo (top) and MU-LLaMA (bottom). The DS branch provides key/value tokens to a gated residual adapter while the original projection provides the query; snowflakes and flames mark frozen and fine-tuned modules. To control for adding capacity or a second pitch-resolved pathway, the CQT and Gaussian conditions reuse the same encoder, fusion point, output dimension, and trainable-parameter budget. They replace D with the aligned normalized magnitude CQT or a fixed random input, respectively. Interpretive scope. DS operationalizes one periodicity/harmonic-distance-inspired relation and its amplitude-weighted localization. It does not explicitly model auditory-filter bandwidths, masking, learned tonal syntax, or cultural preference. The task experiments therefore test whether this representation is useful as an inductive bias, not whether it is a complete perceptual theory of consonance. Experiments Experimental Setup Implementation, compute, and reproducibility. Controlled calculations are deterministic and parameter-free. PyTorch downstream systems run on a Slurm-managed Linux cluster with one NVIDIA A100 per job and no multi-GPU parallelism. CQT and DS features are extracted once per exact excerpt, cached with metadata, and reused across conditions; the submitted environment lock records software and pretrained-model revisions. We use paired seeds 17,42,101,2025,2026,3407\17,42,101,2025,2026,3407\. Within each dataset and seed, all four conditions share splits, excerpts, masks, minibatch order, optimization, early stopping, checkpoint selection, and evaluation; random generators use the listed seed. Except for PMEmo’s released chorus clips, inputs are 24-kHz, 45-second excerpts with zero padding or one deterministic crop shared by all branches. CQT uses a 1,024-sample hop, eight octaves, 72 bins per octave, fmin=C1f_ =C1, and 576 bins; DS applies the proposed transform and log(1+) (1+D). Primary endpoints are BERTScore-R for MusicQA and R2¯VA R^2_VA, the mean of six valence/arousal R2R^2 values, for Music2Emo. Tests use unrounded seed-level values. Controlled Music-Theory Validation We first test whether the relation operator recovers controlled musical structure from rendered piano audio: 13 dyads from unison through the octave, four representative and 13 extended chord voicings, seven diatonic connections to C major, and seven church modes. CQT uses 22.05 kHz audio, four frames per second, eight octaves, 72 bins per octave, fmin=C1f_ =C1, and K=576K=576; the rational search uses Q=60Q=60 and α=.01α=.01. Intervals and scales use a C4 reference, chord qualities equally average tonic-reference and intrinsic DS, and connections use C major. These settings test the same operator under phenomenon-specific reference contexts; downstream models use intrinsic DS. Interval, chord, and connection maps are reduced to Dmax=max∑kt(k,t)D_ = _t _kD(k,t); sequential scales use Dsum=∑t∑k(k,t)D_sum= _t _kD(k,t). Spearman ρ and Kendall τb _b compare these summaries with predefined orders. The interval order broadly agrees with classic dyad studies (Malmberg 1918; Schwartz et al. 2003); chord, function, and mode orders are theory-derived hypotheses, so they test internal music-theoretical consistency rather than population-level perception. Table 1: Controlled validation. Parentheses give two-sided p-values; the four-class row is an ordering check. Test n Spearman ρ Kendall τb _b Intervals 13 .951 (×10−76.36\!×\!10^-7) .846 (×10−65.20\!×\!10^-6) Chord quality (4 exemplars) 4 1.000 (–) 1.000 (–) Chord quality (13 voicings) 13 .626 (.022) .462 (.030) Functional connections 7 .794 (.033) .655 (.054) Church modes 7 .893 (.0068) .810 (.0107) Intervals show the strongest agreement: C–C♯ is maximal, the tritone is high, and the perfect fifth and octave are among the lowest (Figure 5). Audio DS maxima track both the sampled curve and predefined ranks. Four chord exemplars follow major << minor << suspended << diminished; across 13 voicings the positive but nonmonotonic association reflects spacing and inversion rather than chord labels alone. Tonic-function connections are generally below predominant and dominant connections, while Ionian is near the low end and Locrian highest among modes, with local reversals. Item values, spectral maps, loudness sensitivity, timbre, and microtonal analyses are in the supplement. Figure 5: Interval validation: audio DS maxima align with the sampled dissonance curve and the predefined rank reference. Music Question Answering Protocol. MU-LLaMA is fine-tuned on 70,011 question–answer pairs from 7,779 tracks and evaluated on 5,040 pairs from 560 audio-disjoint MTG-Jamendo tracks. All conditions use frozen LLaMA-2 7B and MERT-v1-330M with the released initialization (Touvron et al. 2023; Li et al. 2024; Liu et al. 2024). The Baseline has 4.21M trainable parameters; each matched branch has 5.57M. BF16 AdamW uses gradient accumulation to an effective batch of 32, at most four epochs, and validation-loss early stopping; the best checkpoint is evaluated once on the held-out set. One fixed evaluator reports BLEU, METEOR, ROUGE-L, BERTScore-R, loss, and perplexity (Papineni et al. 2002; Banerjee and Lavie 2005; Lin 2004; Zhang et al. 2020); exact tokenization, generation, and package settings are in the supplement. Table 2: MusicQA results over six paired training seeds (mean± ). Δ is DS minus Baseline. Metric Baseline Gaussian CQT DS Δ BLEU ↑ .2987±.0019 .2985±.0019 .3056±.0015 .3074±.0014 +.0087+.0087 METEOR ↑ .3761±.0017 .3759±.0018 .3838±.0014 .3857±.0013 +.0096+.0096 ROUGE-L ↑ .4556±.0018 .4554±.0019 .4643±.0016 .4671±.0015 +.0115+.0115 BERTScore-R ↑ .8952±.0007 .8950±.0006 .8996±.0006 .9024±.0015 +.0072+.0072 Test loss ↓ .625±.004 .626±.004 .607±.003 .600±.002 −.025-.025 Perplexity ↓ 1.868±.008 1.870±.008 1.836±.005 1.822±.004 −.046-.046 DS has the highest six-seed mean on every MusicQA metric. BERTScore-R increases by .0072.0072 over Baseline, .0074.0074 over Gaussian, and .0028.0028 over CQT, with positive paired differences for all six seeds. The largest Holm-adjusted paired-t p-value is .0017.0017, but the exact two-sided sign test gives .03125.03125 per comparison and .09375.09375 after Holm correction across the three DS-versus-control comparisons. We therefore emphasize repeated-seed direction and effect size, not familywise distribution-free significance. Text metrics measure reference similarity rather than factual musical understanding. Music Emotion Recognition Model and data. Music2Emo projects MERT layers five and six, chord features, and key mode to 512 dimensions before the DS adapter (Kang and Herremans 2025). The Baseline task network has 1.07M trainable parameters; each matched branch adds 0.21M while the 95M-parameter MERT remains frozen. We retain the official MTG-Jamendo split (Bogdanov et al. 2019) and fixed 70/15/15 track splits for DEAM, EmoMusic, and PMEmo (Aljanaki et al. 2017; Soleymani et al. 2013; Zhang et al. 2018; Kang and Herremans 2025). Weighted binary cross-entropy covers 56 tags, and mean squared error covers valence and arousal. All conditions share the knowledge-distillation objective, Adam at 10−410^-4, early stopping, and checkpoint rule; details are in the supplement. Table 3: Music2Emo results over six paired seeds. J-PR/J-ROC are macro tag averages; V/A denote valence/arousal R2R^2. Only R2¯VA R^2_VA reports mean± . Metric Baseline Gaussian CQT DS J-PR .1539.1539 .1537.1537 .1564.1564 .1580 J-ROC .7806.7806 .7801.7801 .7828.7828 .7841 DEAM-V .5169.5169 .5164.5164 .5272.5272 .5355 DEAM-A .6209.6209 .6202.6202 .6260.6260 .6291 Emo-V .6487.6487 .6479.6479 .6575.6575 .6642 Emo-A .7598.7598 .7590.7590 .7642.7642 .7668 PM-V .5451.5451 .5445.5445 .5532.5532 .5587 PM-A .7926.7926 .7920.7920 .7970.7970 .7992 R2¯VA R^2_VA .6473±.0014 .6467±.0014 .6542±.0016 .6589±.0018 DS has the highest mean on every Music2Emo endpoint. R2¯VA R^2_VA increases by .0116.0116 over Baseline, .0122.0122 over Gaussian, and .0047.0047 over CQT, with positive differences for all six seeds. The Holm-adjusted exact sign-test value is .09375.09375, so we again emphasize direction and effect size. Average valence R2R^2 rises from .5702.5702 to .5861.5861, and arousal R2R^2 from .7244.7244 to .7317.7317. CQT is second; the further DS gain is consistent with useful pairwise organization, although compression and normalization differences mean the comparison does not isolate the relation transform alone. Conclusion This work addresses a gap between energy-based spectra and latent learned representations by making one class of simultaneous frequency relations explicit. As a spectrogram reorganizes a waveform without adding a new observation, DS reorganizes magnitude information so that modeled pair relations become directly visible to both researchers and downstream encoders. It applies a continuous, tolerance-based harmonic-distance kernel to magnitude CQT, localizes amplitude-weighted pair relations in time and frequency, and avoids a quadratic pair tensor through frequency-axis correlation. A shape-preserving, zero-initialized adapter then adds this representation to existing systems without changing their output interfaces or baseline function at initialization. Controlled tests recovered strong ordinal agreement for intervals, functional connections, and church modes, with moderate positive agreement across diverse chord voicings. Across six paired seeds, DS also achieved the highest mean on every reported MusicQA and Music2Emo endpoint relative to the baseline, a parameter-matched Gaussian branch, and an architecture-matched magnitude-CQT branch. The Gaussian control indicates that capacity alone is insufficient, while the smaller advantage over magnitude CQT suggests value beyond an added pitch-resolved pathway, within the stated compression and normalization caveat. Together, these results support explicit relational structure as a useful and inspectable complement to learned music representations. Beyond aggregate scores, the retained time–frequency layout gives DS a diagnostic role that scalar dissonance summaries cannot provide. Researchers can inspect when and where a modeled relation is concentrated, while the cross-reference construction can condition that attribution on an explicit tonic, chord, or spectral template. Because extraction is deterministic and cached, and the adapter preserves host shapes and offers exact fallback, the representation can be evaluated alongside existing systems without replacing their audio encoders. These properties make DS a concrete basis for theory-conditioned probing and model comparison, although their practical value outside the evaluated tasks remains to be tested. The evidence does not establish DS as a complete model of consonance or musical preference. The kernel encodes one periodicity/harmonic-distance prior with deliberate octave folding and omits auditory-filter bandwidths, masking, tonal syntax, harsh high-frequency content, rhythmic tension, and culturally or individually learned preference. The controlled rankings are partly theory-derived, no new listening study was conducted, and automatic MusicQA references measure similarity rather than factual or expert-level harmonic reasoning. Evaluation is also limited to two host families and does not fully isolate the relation transform from branch-input compression and normalization. Future work should package the extraction and visualization pipeline as a reusable library, calibrate the representation with listener data and individual perceptual variation, and test its transfer to broader audio and symbolic tasks such as generation, style analysis, and standardized symbolic dissonance encoding. References Aljanaki and Soleymani (2018) A. Aljanaki and M. Soleymani A data-driven approach to mid-level perceptual musical feature modeling. In Proceedings of the 19th International Society for Music Information Retrieval Conference, p. 615–621. Cited by: Introduction, Music representations and perceptual priors., Recorded music and consonance-aware music AI.. Aljanaki et al. (2017) A. Aljanaki, Y. Yang, and M. Soleymani Developing a benchmark for emotional analysis of music. PLOS ONE 12 (3), p. e0173392. External Links: Document Cited by: Model and data.. Banerjee and Lavie (2005) S. Banerjee and A. Lavie METEOR: an automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, p. 65–72. Cited by: Protocol.. Bogdanov et al. (2019) D. Bogdanov, M. Won, P. Tovstogan, A. Porter, and X. Serra The MTG-Jamendo dataset for automatic music tagging. In Machine Learning for Music Discovery Workshop at the 36th International Conference on Machine Learning, Cited by: Model and data.. Brown (1991) J. C. Brown Calculation of a constant-q spectral transform. The Journal of the Acoustical Society of America 89 (1), p. 425–434. External Links: Document Cited by: Music representations and perceptual priors.. Cazden (1980) N. Cazden The definition of consonance and dissonance. International Review of the Aesthetics and Sociology of Music 11 (2), p. 123–168. External Links: Document Cited by: Consonance perception and computational models.. Chowdhury et al. (2019) S. Chowdhury, A. Vall, V. Haunschmid, and G. Widmer Towards explainable music emotion recognition: the route via mid-level features. In Proceedings of the 20th International Society for Music Information Retrieval Conference, p. 237–243. Cited by: Introduction, Music representations and perceptual priors.. Cousineau et al. (2012) M. Cousineau, J. H. McDermott, and I. Peretz The basis of musical consonance as revealed by congenital amusia. Proceedings of the National Academy of Sciences 109 (48), p. 19858–19863. External Links: Document Cited by: Consonance perception and computational models.. Deng et al. (2024) Z. Deng, Y. Ma, Y. Liu, R. Guo, G. Zhang, W. Chen, W. Huang, and E. Benetos MusiLingo: bridging music and text with pre-trained language models for music captioning and query response. In Findings of the Association for Computational Linguistics: NAACL 2024, p. 3643–3655. External Links: Document Cited by: Introduction, Music representations and perceptual priors.. Eerola and Lahdelma (2021) T. Eerola and I. Lahdelma The anatomy of consonance/dissonance: evaluating acoustic and cultural predictors across multiple datasets with chords. Music & Science 4, p. 1–19. External Links: Document Cited by: Consonance perception and computational models.. Elizalde et al. (2023) B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang CLAP: learning audio concepts from natural language supervision. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, p. 1–5. External Links: Document Cited by: Introduction, Music representations and perceptual priors.. Harrison and Pearce (2020) P. M. C. Harrison and M. T. Pearce Simultaneous consonance in music perception and composition. Psychological Review 127 (2), p. 216–244. External Links: Document Cited by: Introduction, Consonance perception and computational models.. Harte et al. (2006) C. Harte, M. Sandler, and M. Gasser Detecting harmonic change in musical audio. In Proceedings of the 1st ACM Workshop on Audio and Music Computing Multimedia, p. 21–26. External Links: Document Cited by: Introduction, Music representations and perceptual priors.. Huang et al. (2022) Q. Huang, A. Jansen, J. Lee, R. Ganti, J. Y. Li, and D. P. W. Ellis MuLan: a joint embedding of music audio and natural language. In Proceedings of the 23rd International Society for Music Information Retrieval Conference, p. 559–566. Cited by: Introduction, Music representations and perceptual priors.. Hutchinson and Knopoff (1978) W. Hutchinson and L. Knopoff The acoustic component of western consonance. Interface 7 (1), p. 1–29. External Links: Document Cited by: Consonance perception and computational models.. Kaklamani and Simserides (2026) S. Kaklamani and C. Simserides Psychoacoustic study of simple-tone dyads: frequency ratio and pitch. Acoustics 8 (1), p. 14. External Links: Document Cited by: Consonance perception and computational models.. Kang and Herremans (2025) J. Kang and D. Herremans Towards unified music emotion recognition across dimensional and categorical models. External Links: 2502.03979, Document Cited by: Model and data., Music2Emo parameters., Table 14. Korzeniowski and Widmer (2016) F. Korzeniowski and G. Widmer Feature learning for chord recognition: the deep chroma extractor. In Proceedings of the 17th International Society for Music Information Retrieval Conference, p. 37–43. Cited by: Introduction, Music representations and perceptual priors.. Langner (1992) G. Langner Periodicity coding in the auditory system. Hearing Research 60 (2), p. 115–142. External Links: Document Cited by: Introduction, Consonance perception and computational models.. Li et al. (2024) Y. Li, R. Yuan, G. Zhang, Y. Ma, X. Chen, H. Yin, C. Xiao, C. Lin, A. Ragni, E. Benetos, N. Gyenge, R. B. Dannenberg, R. Liu, W. Chen, G. Xia, Y. Shi, W. Huang, Z. Wang, Y. Guo, and J. Fu MERT: acoustic music understanding model with large-scale self-supervised training. In International Conference on Learning Representations, External Links: Link Cited by: Introduction, Introduction, Music representations and perceptual priors., Music representations and perceptual priors., Protocol.. Lin (2004) C. Lin ROUGE: a package for automatic evaluation of summaries. In Text Summarization Branches Out, p. 74–81. Cited by: Protocol.. Liu et al. (2024) S. Liu, A. S. Hussain, C. Sun, and Y. Shan Music understanding LLaMA: advancing text-to-music generation with question answering and captioning. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, p. 286–290. External Links: Document Cited by: Introduction, Music representations and perceptual priors., Protocol., Additional MusicQA Comparisons, Table 13. Lu et al. (2021) W. Lu, J. Wang, M. Won, K. Choi, and X. Song SpecTNT: a time–frequency transformer for music audio. In Proceedings of the 22nd International Society for Music Information Retrieval Conference, p. 396–403. Cited by: Introduction, Music representations and perceptual priors.. Malmberg (1918) C. F. Malmberg The perception of consonance and dissonance. Psychological Monographs 25 (2), p. 93–133. External Links: Document Cited by: Controlled Music-Theory Validation, Intervals. Marjieh et al. (2024) R. Marjieh, P. M. C. Harrison, H. Lee, F. Deligiannaki, and N. Jacoby Timbral effects on consonance disentangle psychoacoustic mechanisms and suggest perceptual origins for musical scales. Nature Communications 15, p. 1482. External Links: Document Cited by: Consonance perception and computational models.. McDermott et al. (2016) J. H. McDermott, A. F. Schultz, E. A. Undurraga, and R. A. Godoy Indifference to dissonance in native amazonians reveals cultural variation in music perception. Nature 535 (7613), p. 547–550. External Links: Document Cited by: Consonance perception and computational models.. Müller and Ewert (2011) M. Müller and S. Ewert Chroma toolbox: matlab implementations for extracting variants of chroma-based audio features. In Proceedings of the 12th International Society for Music Information Retrieval Conference, p. 215–220. Cited by: Music representations and perceptual priors.. Panda et al. (2023) R. Panda, R. Malheiro, and R. P. Paiva Audio features for music emotion recognition: a survey. IEEE Transactions on Affective Computing 14 (1), p. 68–88. External Links: Document Cited by: Recorded music and consonance-aware music AI.. Papineni et al. (2002) K. Papineni, S. Roukos, T. Ward, and W. Zhu BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, p. 311–318. External Links: Document Cited by: Protocol.. Parncutt (1989) R. Parncutt Harmony: a psychoacoustical approach. Springer-Verlag, Berlin. External Links: Document Cited by: Consonance perception and computational models.. Plomp and Levelt (1965) R. Plomp and W. J. M. Levelt Tonal consonance and critical bandwidth. The Journal of the Acoustical Society of America 38 (4), p. 548–560. External Links: Document Cited by: Introduction, Consonance perception and computational models.. Poltronieri et al. (2025) A. Poltronieri, X. Serra, and M. Rocamora From discord to harmony: decomposed consonance-based training for improved audio chord estimation. In Proceedings of the 26th International Society for Music Information Retrieval Conference, Daejeon, South Korea, p. 492–502. External Links: Document Cited by: Introduction, Music representations and perceptual priors.. Schwär et al. (2025) S. Schwär, S. Balke, and M. Müller Measuring sensory dissonance in multi-track music recordings: a case study with wind quartets. In Proceedings of the 26th International Society for Music Information Retrieval Conference, Daejeon, South Korea, p. 117–126. External Links: Document Cited by: Recorded music and consonance-aware music AI.. Schwartz et al. (2003) D. A. Schwartz, C. Q. Howe, and D. Purves The statistical structure of human speech sounds predicts musical universals. The Journal of Neuroscience 23 (18), p. 7160–7168. External Links: Document Cited by: Controlled Music-Theory Validation, Intervals. Sethares (1994) W. A. Sethares Adaptive tunings for musical scales. The Journal of the Acoustical Society of America 96 (1), p. 10–18. Cited by: Amplitude-weighted attribution.. Sethares (2005) W. A. Sethares Tuning, timbre, spectrum, scale. 2 edition, Springer, London. External Links: Document Cited by: Consonance perception and computational models., Amplitude-weighted attribution.. Soleymani et al. (2013) M. Soleymani, M. N. Caro, E. M. Schmidt, C. Sha, and Y. Yang 1000 songs for emotional analysis of music. In Proceedings of the 2nd ACM International Workshop on Crowdsourcing for Multimedia, p. 1–6. External Links: Document Cited by: Model and data.. Stolzenburg (2015) F. Stolzenburg Harmony perception by periodicity detection. Journal of Mathematics and Music 9 (3), p. 215–238. External Links: Document Cited by: Introduction, Introduction, Consonance perception and computational models., Overview, Rational candidates and complexity.. Tenney (1984) J. Tenney John cage and the theory of harmony. In Soundings 13: The Music of James Tenney, P. Garland (Ed.), p. 55–83. Cited by: Introduction, Overview, Rational candidates and complexity.. Touvron et al. (2023) H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama 2: open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. External Links: Document Cited by: Protocol.. von Helmholtz (1954) H. L. F. von Helmholtz On the sensations of tone as a physiological basis for the theory of music. Dover Publications, New York. Note: Second English edition, translated by Alexander J. Ellis; original German work published in 1863 Cited by: Introduction, Consonance perception and computational models.. Wagner et al. (2022) B. Wagner, S. Czoschke, B. Tillmann, S. Koelsch, P. Vuust, and E. Brattico Pitch chroma information is processed in addition to pitch height information with more than two pitch-range categories. Attention, Perception, & Psychophysics 84 (5), p. 1757–1771. External Links: Document Cited by: Continuous pitch intervals and octave folding.. Wand (2012) A. Wand On the conception and measure of consonance. Leonardo Music Journal 22, p. 73–78. External Links: Document Cited by: Consonance perception and computational models.. Weck et al. (2026) B. Weck, P. Puentes, A. Poltronieri, S. Prabhu, and D. Bogdanov HumMusQA: a human-written music understanding QA benchmark dataset. In Proceedings of the 4th Workshop on NLP for Music and Audio (NLP4MusA 2026), Rabat, Morocco, p. 58–67. External Links: Document Cited by: Music representations and perceptual priors.. Won et al. (2024) M. Won, Y. Hung, and D. Le A foundation model for music informatics. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, p. 1226–1230. External Links: Document Cited by: Introduction, Music representations and perceptual priors.. Xia et al. (2026) H. Xia, Z. Huang, Y. Tan, and S. Song Let the model learn to feel: mode-guided tonality injection for symbolic music emotion recognition. Proceedings of the AAAI Conference on Artificial Intelligence 40 (3), p. 2182–2190. External Links: Document Cited by: Introduction. Zhang et al. (2018) K. Zhang, H. Zhang, S. Li, C. Yang, and L. Sun The PMEmo dataset for music emotion recognition. In Proceedings of the 2018 ACM International Conference on Multimedia Retrieval, p. 135–142. External Links: Document Cited by: Model and data.. Zhang et al. (2020) T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi BERTScore: evaluating text generation with BERT. In International Conference on Learning Representations, Cited by: Protocol.. Supplementary Material: Beyond Frequency: Dissonance Spectrum for Perceptually Motivated Music Understanding 1Beijing Institute for General Artificial Intelligence, Beijing, China 2Central Conservatory of Music, Beijing, China 3Peking University, Beijing, China Reproducibility Package The separately submitted Code and Data Archive contains the DS implementation, extraction configuration, cached-feature metadata schema, fixed dataset manifests, audio hashes, training configurations, environment lock files, checkpoint-selection rules, evaluation scripts, seed-level predictions, and scripts that regenerate every reported table. It includes scripts and instructions for obtaining public datasets and pretrained models from their cited official sources. The archive is packaged without author-identifying repository metadata. Full Derivation and Extensions of Dissonance Spectrum This section records the complete algebra and optional variants underlying the concise main-paper description and clarifies implementation shapes and experimental conventions. Preprocessing Conventions Let (k,t)X(k,t) denote the magnitude CQT, x→t=(:,t)∈ℝK x^\,t=X(:,t) ^K, and xkt=(k,t)x_k^\,t=X(k,t). The denoising and peak-selection operators are denoised(k,t)=(k,t)⋅((k,t)≥θ),X_denoised(k,t)=X(k,t)·1 (X(k,t)≥θ ), (1) x→peakt=x→t⊙[(xkt is a peak point in x→t)]k=0K−1. x^\,t_peak= x^\,t [1 (x_k^\,t is a peak point in x^\,t ) ]_k=0^K-1. (2) Two normalization choices are useful in different settings: m(t) m(t) =maxkxkt, = _kx_k^\,t, x→framet x^\,t_frame =x→t/m(t), = x^\,t/m(t), (3) m m =maxk,txkt, = _k,tx_k^\,t, global _global =/m. =X/m. The downstream model comparisons use excerpt-level global-maximum normalization m, matching the magnitude-CQT control and avoiding framewise loudness equalization. The controlled music-theory tests instead use the context-window normalization described below, with target and reference spectra normalized separately. Each comparison uses one fixed convention recorded in its configuration. Direct Pairwise Form The intrinsic element is dkt=1K∑l=0K−1xktxltD¯(pl−pk),d^t_k= 1K _l=0^K-1x^t_k\,x^t_l\, D(p_l-p_k), (4) and the complete representation is =[d→ 0d→ 1⋯d→T−1]∈ℝK×T.D= bmatrix d^\,0& d^\,1&·s& d^\,T-1 bmatrix ^K× T. (5) The factor 1/K1/K averages the reference-bin contributions and removes their direct linear count factor under otherwise matched grids; it is not a guarantee of invariance to CQT resolution. The direct pair-attribution concept is visualized in the main paper. Index Derivation of the Correlation Form The CQT pitch grid and the zero-indexed stored relation vector are pk=pitchmin+12kB,k=0,…,K−1,p_k=pitch_ + 12kB, k=0,…,K-1, (6) [±→]j=D¯(12(j−K+1)B),j=0,…,2K−2. [ D^± ]_j= D\! ( 12(j-K+1)B ), j=0,…,2K-2. (7) Because the kernel depends only on pitch difference, D¯(pl−pk)=±→K+l−k−1. D(p_l-p_k)= D^±_K+l-k-1. (8) For a length-(2K−1)(2K-1) vector a→ a and a length-K vector b→ b, the implementation uses the target-aligned convention [a→⋆valid,fb→]k=∑l=0K−1aK+l−k−1bl,k=0,…,K−1.[ a _valid,f b]_k= _l=0^K-1a_K+l-k-1b_l, k=0,…,K-1. (9) With increasing-lag storage, this equals a standard valid cross-correlation followed by a fixed frequency-axis reversal; an equivalent lag-reversed layout avoids the explicit reversal. Consequently, dkt d^\,t_k =1Kxkt∑l=0K−1xltD¯(pl−pk) = 1Kx_k^t _l=0^K-1x^t_l\, D(p_l-p_k) (10) =1Kxkt∑l=0K−1xltK+l−k−1±. = 1Kx_k^t _l=0^K-1x_l^t\,D^±_K+l-k-1. (11) Sliding the complete relation vector over the CQT gives the reference-relation frame r→t=±→⋆valid,fx→t, r^\,t= D^± _valid,f x^\,t, (12) and therefore d→t=1Kr→t⊙x→t=1K(±→⋆valid,fx→t)⊙x→t. d^\,t= 1K r^\,t x^t= 1K ( D^± _valid,f x^\,t ) x^t. (13) Figure 1 visualizes the index alignment. Equation 9 is applied independently at every time frame and returns the K target locations in their original frequency order. Figure 1: Index alignment for the frequency-axis correlation form of intrinsic DS. Figure 2: Vectorized intrinsic-DS computation after replacing the explicit pairwise sum by valid correlation. Cross Dissonance Spectrum For target x→ x and an arbitrary reference spectrum x→(ref) x^(ref), the cross-spectrum is dkx→x→(ref) d_k^\, x→ x^(ref) =1K∑l=0K−1xkxl(ref)D¯(pl−pk) = 1K _l=0^K-1x_kx_l^(ref)\, D(p_l-p_k) (14) =1Kxk∑l=0K−1xl(ref)K+l−k−1±. = 1Kx_k _l=0^K-1x_l^(ref)\,D^±_K+l-k-1. (15) Its reusable reference map and final attribution are r→=±→⋆valid,fx→(ref), r= D^± _valid,f x^(ref), (16) d→x→x→(ref)=1K(±→⋆valid,fx→(ref))⊙x→. d^\, x→ x^(ref)= 1K ( D^± _valid,f x^(ref) ) x. (17) Intrinsic DS is the special case x→(ref)=x→ x^(ref)= x. A fixed reference may instead encode a tonic template or an instrument tone. Figure 3 shows the corresponding cross-reference computation. Figure 3: Index alignment for cross-DS correlation using a fixed or time-varying reference spectrum. Figure 4: Vectorized cross-DS computation using a reusable reference-relation vector. Matrix and Temporal Variants Repeating the same relation vector across time gives ±=[→±→±⋯→±]∈ℝ(2K−1)×T. D^±= bmatrix D^±& D^±&·s& D^± bmatrix ^(2K-1)× T. (18) For a general reference CQT, =±⋆valid,f(ref),R= D^± _valid,fX^(ref), (19) =1K⊙=1K(±⋆valid,f(ref))⊙.D= 1K\,R = 1K\,( D^± _valid,fX^(ref)) . (20) Both R and D have shape K×TK× T under the frequency-axis valid-correlation convention used here. Optional temporal DS averages references from the preceding n frames: d→(temporal),t=1n∑i=0n−1d→(x→t→x→t−i)=d→(x→t→1n∑i=0n−1x→t−i). d^\,(temporal),\,t= 1n _i=0^n-1 d ( x^\,t→ x^\,t-i )= d ( x^\,t→ 1n _i=0^n-1 x^\,t-i ). (21) A tonic–intrinsic combination is d→(combine),t=wintrinsicd→t+wtonicd→(tonic),t. d^\,(combine),\,t=w_intrinsic d^\,t+w_tonic d^\,(tonic),\,t. (22) The downstream model comparisons use intrinsic DS. The controlled music-theory tests use tonic, chord, or tonic–intrinsic references to isolate the relation being examined. Controlled Music-Theory Validation Details Stimuli, References, and Aggregation The validation suite uses MIDI-rendered stimuli with fixed note events and piano timbre unless otherwise stated. Nominal durations describe MIDI events; WAV files include onset and release tails. The controlled tests are not a listener study. They ask whether the relation operator and its audio realization reproduce predefined ordinal structures under transparent reference choices. Table 1: Controlled-validation configuration. Item Setting Audio / frame rate 22.05 kHz mono / 4 frames s-1 CQT grid 8 octaves, 72 bins octave-1, K=576K=576, fmin=C1f_ =C1 Kernel Q=60Q=60, relative tolerance α=.01α=.01, octave folding Preprocessing linear magnitude, relative-frame soft gate at 0.1, context window 4 s Peak selection prominence .015; other peak filters disabled Scaling target and reference normalized separately; final factor 1/K1/K Intervals / scales C4 tonic reference Chord qualities equal-weight C4-tonic and intrinsic DS Chord connections C-major chord reference Aggregation DmaxD_ for intervals/chords/connections; DsumD_sum for scales For a frame, D(t)=∑k(k,t)D(t)= _kD(k,t). We use Dmax=maxtD(t)D_ = _tD(t) when a stimulus contains a sustained interval or chord event and Dsum=∑tD(t)D_sum= _tD(t) when an entire sequential scale is the analysis unit. Rank hypotheses are min–max mapped only for visualization; all reported correlations use the original ranks and unscaled DS summaries. Figure 5: Four-octave view of the normalized relation kernel, illustrating its octave-folded repetition. Integer-semitone samples are marked for reference. Intervals Thirteen notes from C4 to C5 are compared against a C4 reference. The historical order used by the validation assets is related to classic dyad-consonance experiments (Malmberg 1918; Schwartz et al. 2003), but it is not presented as a direct transcription of any single published table. Table 2: Interval results under the C4 tonic reference. Note Rank DmaxD_ Note Rank DmaxD_ C4 1 .000253 G4 3 .001145 C♯ 4 13 .004379 G♯ 4 7 .001913 D4 11 .003476 A4 6 .001703 D♯ 4 8 .002463 A♯ 4 10 .002583 E4 5 .002298 B4 12 .002802 F4 4 .001614 C5 2 .000030 F♯ 4 9 .003391 Spearman ρ=.95055ρ=.95055 (p=6.36×10−7p=6.36× 10^-7) and Kendall τb=.84615 _b=.84615 (p=5.20×10−6p=5.20× 10^-6). The minor second is maximal, the tritone is high, the perfect fifth is low, and the octave is minimal. The small nonzero unison value reflects slight differences between independently rendered target and reference recordings. Figure 6: Interval validation details: audio DS maxima, sampled kernel values, and the predefined rank reference. Figure 7: Representative interval DS maps for unison, minor second, perfect fifth, and octave. Chord Qualities The four representative classes are perfectly ordered as major << minor << suspended << diminished. Because n=4n=4, this is treated as an ordering check rather than a meaningful significance test. The larger set varies chord quality and voicing and therefore provides a stricter test. Table 3: Extended chord-quality validation. Semitone arrays are relative to C4. Chord Semitones Rank DmaxD_ Chord Semitones Rank DmaxD_ Major-1 [0,4,7] 1 .005691 Sus-1 [0,7,17] 7 .005641 Major-2 [0,3,8] 5 .006941 Sus-2 [0,2,7] 6 .006758 Major-3 [0,9,17] 3 .006299 Sus-3 [0,10,17] 4 .007769 Minor-1 [0,3,7] 2 .006058 Dim-1 [0,3,6] 12 .015950 Minor-2 [0,4,9] 10 .006738 Dim-2 [0,9,15] 9 .008581 Minor-3 [0,8,17] 8 .006363 Dim-3 [0,9,18] 11 .008172 Aug [0,4,8] 13 .007305 Across 13 voicings, Spearman ρ=.62637ρ=.62637 (p=.02199p=.02199) and Kendall τb=.46154 _b=.46154 (p=.03048p=.03048). The positive association coexists with local reversals, such as Sus-1 below Major-1 and the augmented triad below Dim-1. These deviations are informative rather than errors: DS analyzes realized spectra and therefore responds to spacing and inversion instead of assigning one constant to a chord label. Figure 8: Representative chord-quality ordering. Figure 9: Extended chord-quality and voicing results. Figure 10: Representative tonic-referenced chord-quality spectra. Functional Chord Connections Each stimulus plays one diatonic chord followed by C major. The ordinal code assigns 1 to tonic-function chords (I, i, vi), 2 to predominant chords (i, IV), and 3 to dominant-function chords (V7, vii∘). This reference-conditioned construction operationalizes a connection as the first chord’s relation to C major; it does not model voice leading or learned temporal expectation. The code is a theory-derived trend hypothesis, not a psychophysical scale. Table 4: Chord connections to C major. File First chord Group Rank DmaxD_ C C–E–G tonic 1 .003848 DC D–F–A predominant 2 .009701 EC E–G–B tonic 1 .005311 FC F–A–C predominant 2 .008495 G7C G–B–D–F dominant 3 .010312 AC A–C–E tonic 1 .005695 BC B–D–F dominant 3 .006699 Spearman ρ=.79373ρ=.79373 (p=.03310p=.03310) and Kendall τb=.65465 _b=.65465 (p=.05363p=.05363). The overall group trend is recovered, although vii∘ and V7 differ substantially, showing that group membership does not determine the complete spectral value. Figure 11: Functional chord-connection results. Figure 12: Reference-conditioned spectra for the seven chord connections. Scales and Modes Scale stimuli ascend from C4 with one second per note. For the seven church modes, DsumD_sum correlates with the predefined order at Spearman ρ=.89286ρ=.89286 (p=.00681p=.00681) and Kendall τb=.80952 _b=.80952 (p=.0107p=.0107). Table 5: Church-mode results. Mode Rank DsumD_sum Ionian 1 .030220 Mixolydian 2 .030273 Lydian 3 .033697 Dorian 4 .030704 Aeolian 5 .031444 Phrygian 6 .033937 Locrian 7 .038779 The trend is strong but not strictly monotonic because Dorian lies below Lydian. This test aggregates tonic-relative interval content and should not be read as a complete model of modal perception. Figure 14 includes all 33 available scale files. For scales without an externally grounded rank, the values should be interpreted as a quantitative “consonance palette”: DS can compare and visualize their realized interval content, but the ordering is not claimed as a universal preference scale. The code defines a natural-minor file that is absent; the equivalent Aeolian pitch-class set is available. Figure 13: Church-mode total DS and the predefined ordinal reference used in the main-paper correlation. Figure 14: Circular peak patterns for all available scale recordings. Only the seven church modes are used for the ordinal correlation in the main paper. Figure 15: DS curves for the available scale recordings, illustrating scale-dependent consonance profiles. Timbre and Microtonal Examples The same C4–C5 chromatic sequence yields different maxima across seven software instruments, demonstrating sensitivity to partial amplitudes, envelopes, and noise. Because loudness, spectral centroid, and envelope are not matched, these values are descriptive comparisons of complete rendered sounds rather than a causal timbre ranking. A 24-tone-equal-temperament sequence further demonstrates that the continuous kernel can analyze 50-cent steps; no experiential ranking is imposed. Table 6: Exploratory timbre and microtonal examples. Audio Duration (s) DmaxD_ Audio Duration (s) DmaxD_ Saxophone 14.005 .003134 Violins 15.531 .005633 Guitar 14.005 .003170 Alto 15.214 .007333 Piano 14.099 .004385 Trumpet 14.005 .011913 Flute 15.724 .029903 24-TET sequence 26.005 .003751 Figure 16: Exploratory 24-TET spectra at 50-cent resolution. The timbre values in Table 6 are descriptive and are not assigned a universal ranking. Loudness Disentanglement under the Normalized Pipeline The primary quantitative test uses seven independently rendered C4–C♯ 4 dyads spanning the configured level conditions (n=7n=7). Let Li=∑n|yi[n]|L_i= _n|y_i[n]| be the absolute-amplitude level proxy used by the validation suite (not a perceptual loudness measure), and let YiY_i be the DS peak. For positive LiL_i and YiY_i, we fit the robust log–log relation lnYi Y_i =a+βlnLi+ϵi, =a+β L_i+ _i, (23) Delastic D_elastic =max(0,1−|β|), = (0,1-|β|), (24) where β is the Theil–Sen slope and a larger DelasticD_elastic indicates lower proportional sensitivity. The estimate is β=−.000β=-.000 with a Theil–Sen 95% CI of [−.014,.000][-.014,.000], giving Delastic=1.000D_elastic=1.000 after rounding. The DS peak varies by only .94% relative standard deviation and 2.49% relative range, defined as sY/Y¯s_Y/ Y and (maxiYi−miniYi)/Y¯( _iY_i- _iY_i)/ Y, respectively. Thus even the lower confidence bound corresponds to approximately a .014% DS change per 1% loudness change. Pearson r=−.744r=-.744 (p=.0551p=.0551) is also reported, but it describes the ordering of small residual deviations rather than their proportional magnitude; a sizeable |r||r| can therefore coexist with near-zero elasticity. Because the controlled target and reference contexts are normalized, this sanity check supports near gain-invariance of the normalized DS summary over the tested range rather than universal statistical independence from perceptual loudness. The companion crescendo stimulus visualizes the time-resolved behavior within one file. It is not included in the seven-sample elasticity estimate because adjacent events share a rendering and temporal context. Future tests should expand the number of independent levels and instruments, report LUFS and clipping diagnostics, and repeat rendering under multiple gain and normalization conventions. Figure 17: Primary seven-level result: the absolute-amplitude level proxy changes substantially while the maximum DS peak remains tightly concentrated. Figure 18: Level-controlled stimuli. Left: CQT magnitude. Right: intrinsic DS. Figure 19: Within-file crescendo visualization. Left: CQT magnitude. Right: intrinsic DS. Detailed Configuration Compute and software. All downstream runs use one NVIDIA A100 GPU per Slurm job on Linux; no multi-GPU parallelism is used. The implementations are in PyTorch. CQT and DS extraction is cached before training, and the environment lock file in the code archive records exact framework, CUDA, evaluation-package, and pretrained-model revisions. Shared preprocessing. Except for PMEmo’s released chorus clips, all branches use 24-kHz, 45-second excerpts. Magnitude CQT uses a 1,024-sample hop, eight octaves, 72 bins per octave, fmin=C1f_ =C1, and excerpt-level global-maximum normalization. DS is nonnegative and uses log(1+) (1+D) compression after the relation transform. CQT and DS use the same 576-bin grid and identical 576-to-144 pooling, temporal masks, encoder, fusion location, parameter count, optimizer, schedule, and checkpoint selection. The comparison therefore matches architecture and spectral grid but does not isolate the relation transform from branch-input compression and normalization. Fixed Gaussian inputs are generated once per excerpt and training seed and reused across epochs. MU-LLaMA parameters. The Baseline contains 4,205,568 trainable parameters. The matched parallel branch adds 1,360,193 parameters, yielding 5,565,761 trainable parameters in Gaussian, CQT, and DS. The branch uses three convolutional blocks, a 128-dimensional temporal encoder, a kernel-3 depthwise temporal convolution, single-head temporal attention, and gated residual fusion. The output projection is zero-initialized and the gate bias is −3-3. Music2Emo parameters. The Baseline task network contains 1,071,617 trainable parameters. The parallel encoder and gated cross-attention residual module add 209,345 parameters, yielding 1,280,962 parameters in Gaussian, CQT, and DS. The added branch is 0.22% of the frozen 95M-parameter MERT encoder and 19.5% of the trainable task network. Following the Music2Emo 70/15/15 protocol (Kang and Herremans 2025), we generate the track-level split once and keep it fixed across all conditions: 1,261/271/270 for DEAM, 495/124/125 for EmoMusic, and 536/116/115 for the 767-track labeled PMEmo subset. These are experiment-specific splits, not official dataset splits. PMEmo contains 794 items in total and uses the released chorus clip. Evaluation Implementation MusicQA uses the archive’s explicitly specified evaluator, which differs from the released MU-LLaMA scoring script. NLTK word-punctuation tokenization is applied before sentence BLEU with weights (.25,.25,.25,.25)(.25,.25,.25,.25) and method-1 smoothing. METEOR uses exact, stem, and WordNet synonym matching. ROUGE-L uses rouge-score with stemming and reports F-measure. BERTScore uses roberta-large in English without IDF or baseline rescaling and reports recall. Loss covers answer tokens; perplexity is exponentiated per seed before averaging. For MTG-Jamendo, PR-AUC and ROC-AUC are macro-averaged over the 56 tags. The Music2Emo weighted binary cross-entropy uses wi=2/(1+pi)w_i=2/(1+p_i) for the positive term and w¯i=2pi/(1+pi) w_i=2p_i/(1+p_i) for the negative term, where pip_i is the training prevalence of tag i. The lock file records exact package and model revisions. Additional MusicQA Comparisons The MusicQA references and model outputs are free-form text. To avoid cherry-picking fluent examples, we summarize additional corpus-level comparisons rather than selecting answers by visual inspection. The questions follow the MU-LLaMA data-generation setting and cover attributes such as mood, instrumentation, tempo, genre, and overall tone (Liu et al. 2024). Table 7 reports the difference between the DS condition and each matched control using the six-seed means from the fixed 5,040-pair evaluation set. Higher is better for text-similarity metrics and lower is better for loss and perplexity. Seed-level predictions, questions, and references are included in the separately submitted archive for answer-level inspection. Table 7: Additional MusicQA comparisons. Entries are DS minus the indicated control; negative values are improvements for loss and perplexity. Metric vs. Baseline vs. Gaussian vs. CQT BLEU ↑ +.0087 +.0089 +.0018 METEOR ↑ +.0096 +.0098 +.0019 ROUGE-L ↑ +.0115 +.0117 +.0028 BERTScore-R ↑ +.0072 +.0074 +.0028 Test loss ↓ -.025 -.026 -.007 Perplexity ↓ -.046 -.048 -.014 The Gaussian comparison shows that the improvement is not explained by the added parameter budget alone. The smaller but consistently favorable mean difference over the architecture-matched CQT branch indicates that the pitch-resolved input accounts for part, but not all, of the gain. These comparisons remain reference-similarity analyses: automatically generated references can reward paraphrase overlap and do not by themselves establish factual correctness or expert-level harmonic reasoning. We therefore treat the qualitative prediction files as audit material and the paired BERTScore-R endpoint as the prespecified aggregate comparison. Answer-level audit examples. Tables 8–10 reproduce all ten rows in the supplied audit set rather than selecting a fluent subset. Table 8 maps each case to the audio identifier and MTG-Jamendo track index recorded by the accompanying demo. The Original Token F1 column is the token-overlap score between the original MU-LLaMA output and the reference answer; it is not a score for the DS-augmented output. Tables 9 and 10 provide the corresponding questions, reference answers, and both inference outputs. Relative to the original outputs, the DS-augmented outputs tend to replace generic genre labels, inferred lyric narratives, or vague “weird” descriptions with more specific acoustic, stylistic, and affective attributes. These examples are qualitative audit material and do not replace the aggregate paired evaluation. Table 8: Provenance and score ownership for the MusicQA audit examples. Audio ID and dataset track index reproduce the audio and trackId fields in the accompanying demo. Original Token F1 applies only to the original MU-LLaMA output. ID Audio ID Dataset track index Original Token F1 ID Audio ID Dataset track index Original Token F1 01 012895.wav 661609 .000 06 013055.wav 873804 .071 02 012815.wav 477309 .038 07 012553.wav 8308 .111 03 012725.wav 267904 .063 08 012986.wav 796003 .073 04 012786.wav 387501 .087 09 012972.wav 768506 .094 05 012750.wav 345903 .000 10 013063.wav 875005 .103 Table 9: MusicQA answer-level audit examples, cases 01–05. Reference answers and dataset identifiers are reproduced from the accompanying demo. ID Question Reference answer Original MU-LLaMA DS-augmented MU-LLaMA 01 What is the term used to describe a style of music that features a strong emphasis on melody and harmony? Acoustic rock is a term used to describe a style of music that features a strong emphasis on melody and harmony. Acoustic rock typically combines elements of rock and folk music, with acoustic guitar, harmonica, and vocals as the main instruments. Classical. Acoustic rock. 02 What type of orchestral music is characterized by complex harmonies and intricate compositions? Progressive orchestral. Progressive orchestral music incorporates elements of progressive rock, avant-garde, and experimental music, and is characterized by complex harmonies and intricate compositions. Symphonic music. Progressive orchestral music. 03 What can be inferred from the audio? The audio blends folk, punk rock, rock, and ska. It uses acoustic guitars, bass, drums, and occasional brass and strings, with gruff vocals, harmonies, and call-and-response sections. The song is about a man’s love for a woman. The music blends folk, punk rock, ska, and acoustic instrumentation. 04 Describe the audio in detail. The audio is a soothing and upbeat tune with a catchy melody and harmonious vocals. Harp, synthesizer, and piano create an elegant, layered, and enjoyable arrangement. The music is described as slow and weird. The music is described as soothing, upbeat, and harmonious. 05 What is the genre of music that is characterized by a slow and smooth rhythm, often played on a saxophone? Bossanova. It originated in Brazil and is known for its relaxed and romantic sound, often featuring a slow and smooth rhythm. Jazz. Bossanova. Table 10: MusicQA answer-level audit examples, cases 06–10 (continued). ID Question Reference answer Original MU-LLaMA DS-augmented MU-LLaMA 06 What can be inferred from the audio? The audio appears to be a movie or TV soundtrack with a moody, atmospheric character. Its slow, mellow melody and layered guitars and piano evoke longing, nostalgia, and introspection. The song is about love and heartbreak. The music feels moody, mellow, nostalgic, and introspective. 07 Describe the audio in detail. A relaxing instrumental composition combining electronic, lounge, and oriental music. Its slow, smooth tempo and minimal instrumentation create a calming atmosphere. The music is described as having a weird and otherworldly quality. The music is described as relaxing, smooth, and gently oriental. 08 What can be inferred from the audio? A relaxing and upbeat trip-hop track for chilling out. Synthesizer, strings, bass, drums, Rhodes, and trumpet create a dreamy, smooth, and sophisticated atmosphere. The song is about love and heartbreak. The track is relaxing, upbeat, dreamy, and sophisticated. 09 Describe the audio in detail. A calm orchestral composition. Acoustic bass, cello, piano, violin, bell, and oboe create a warm, elegant, magical, and serene piece for relaxing. The music is described as having a weird and otherworldly quality. The music is described as calm, warm, elegant, and serene. 10 Describe the audio in detail. A fusion of classical and ambient music with flute and synthesizer, a slow and mellow pace, bass and electric guitar foundation, and soulful background vocals. The music is described as having a weird and scary atmosphere. The music is described as slow, mellow, soulful, and ambient. Exploratory Mechanism Analysis The mechanism analyses are secondary and are not used to select the reported architecture. MusicQA uses seeds 17,42,101\17,42,101\; Music2Emo uses all six paired seeds. Global pooling removes temporal localization, temporal shuffling preserves marginal token statistics while destroying order, pre-projection or early fusion changes the insertion point, and the randomized kernel preserves symmetry and its value distribution while permuting frequency correspondence. Table 11: Exploratory MusicQA mechanism analysis. Variant BERTScore-R ↑ Δ vs. proposed Global DS .8985±.0007 −.0030-.0030 Randomized kernel .8992±.0009 −.0023-.0023 Temporal DS, shuffled .8997±.0007 −.0018-.0018 Temporal DS, pre-projection .9004±.0006 −.0011-.0011 Temporal DS, proposed .9015±.0013 – Table 12: Exploratory Music2Emo mechanism analysis. Variant Emotion R2¯VA R^2_VA ↑ Global DS .6524±.0015 Temporal DS, shuffled .6550±.0017 Global DS, early concatenation .6558±.0016 Temporal DS, proposed .6589±.0018 All perturbations reduce the mean relative to the proposed ordered temporal configuration. These results are consistent with a contribution from frequency correspondence and temporal organization, but they do not isolate a single causal mechanism and support no inferential claim. Reference Comparisons under Modified Protocols Table 13: Contextual MU-LLaMA comparison. The original values use the released scoring script on 4,500 QA pairs from 500 tracks (Liu et al. 2024); our values use the explicitly specified evaluator and are six-seed means on the custom 5,040-pair, 560-track set. Source BLEU METEOR ROUGE-L BERTScore-R Original report (released script) .3060 .3850 .4660 .9010 Our baseline (specified metrics) .2987 .3761 .4556 .8952 Absolute difference .0073 .0089 .0104 .0058 The rows are not treated as a direct reproduction comparison because both the evaluation set and metric implementation differ. Table 14: Contextual Music2Emo comparison. The original report uses its 30-second segment augmentation and original training protocol (Kang and Herremans 2025); our baseline uses the modified fixed-excerpt protocol described here. Source J-PR J-ROC DEAM-V DEAM-A Emo-V Emo-A PM-V PM-A Original report .1543 .7810 .5184 .6228 .6512 .7616 .5473 .7940 Our baseline .1539 .7806 .5169 .6209 .6487 .7598 .5451 .7926 Absolute difference .0004 .0004 .0015 .0019 .0025 .0018 .0022 .0014 Seed-Level Primary Endpoints Table 15: MusicQA BERTScore-R for the six paired seeds. Seed Baseline Gaussian CQT DS 17 .8943 .8942 .8990 .9001 42 .8950 .8948 .8996 .9017 101 .8954 .8951 .8998 .9027 2025 .8948 .8947 .8989 .9028 2026 .8962 .8960 .8997 .9024 3407 .8955 .8952 .9006 .9047 Mean .8952 .8950 .8996 .9024 SD .0007 .0006 .0006 .0015 Table 16: Music2Emo R2¯VA R^2_VA for the six paired seeds. Seed Baseline Gaussian CQT DS 17 .6453 .6449 .6519 .6560 42 .6472 .6466 .6542 .6587 101 .6481 .6476 .6552 .6596 2025 .6464 .6457 .6525 .6580 2026 .6494 .6488 .6557 .6602 3407 .6476 .6468 .6555 .6609 Mean .6473 .6467 .6542 .6589 SD .0014 .0014 .0016 .0018 Paired Comparisons Table 17: Paired comparisons on the two designated primary endpoints using unrounded seed-level values. pH(t)p_H^(t) is Holm-adjusted paired-t; pH(sign)p_H^(sign) is the Holm-adjusted exact two-sided sign test based only on paired directions. Each task has one three-comparison Holm family. Task Comparison Mean Δ pH(t)p_H^(t) pH(sign)p_H^(sign) MusicQA DS vs. Baseline +.0072 <.001<.001 .0938 MusicQA DS vs. Gaussian +.0074 <.001<.001 .0938 MusicQA DS vs. CQT +.0028 .0017 .0938 Music2Emo DS vs. Baseline +.0116 <.001<.001 .0938 Music2Emo DS vs. Gaussian +.0122 <.001<.001 .0938 Music2Emo DS vs. CQT +.0047 <.001<.001 .0938 References Aljanaki and Soleymani (2018) A. Aljanaki and M. Soleymani A data-driven approach to mid-level perceptual musical feature modeling. In Proceedings of the 19th International Society for Music Information Retrieval Conference, p. 615–621. Cited by: Introduction, Music representations and perceptual priors., Recorded music and consonance-aware music AI.. Aljanaki et al. (2017) A. Aljanaki, Y. Yang, and M. Soleymani Developing a benchmark for emotional analysis of music. PLOS ONE 12 (3), p. e0173392. External Links: Document Cited by: Model and data.. Banerjee and Lavie (2005) S. Banerjee and A. Lavie METEOR: an automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, p. 65–72. Cited by: Protocol.. Bogdanov et al. (2019) D. Bogdanov, M. Won, P. Tovstogan, A. Porter, and X. Serra The MTG-Jamendo dataset for automatic music tagging. In Machine Learning for Music Discovery Workshop at the 36th International Conference on Machine Learning, Cited by: Model and data.. Brown (1991) J. C. Brown Calculation of a constant-q spectral transform. The Journal of the Acoustical Society of America 89 (1), p. 425–434. External Links: Document Cited by: Music representations and perceptual priors.. Cazden (1980) N. Cazden The definition of consonance and dissonance. International Review of the Aesthetics and Sociology of Music 11 (2), p. 123–168. External Links: Document Cited by: Consonance perception and computational models.. Chowdhury et al. (2019) S. Chowdhury, A. Vall, V. Haunschmid, and G. Widmer Towards explainable music emotion recognition: the route via mid-level features. In Proceedings of the 20th International Society for Music Information Retrieval Conference, p. 237–243. Cited by: Introduction, Music representations and perceptual priors.. Cousineau et al. (2012) M. Cousineau, J. H. McDermott, and I. Peretz The basis of musical consonance as revealed by congenital amusia. Proceedings of the National Academy of Sciences 109 (48), p. 19858–19863. External Links: Document Cited by: Consonance perception and computational models.. Deng et al. (2024) Z. Deng, Y. Ma, Y. Liu, R. Guo, G. Zhang, W. Chen, W. Huang, and E. Benetos MusiLingo: bridging music and text with pre-trained language models for music captioning and query response. In Findings of the Association for Computational Linguistics: NAACL 2024, p. 3643–3655. External Links: Document Cited by: Introduction, Music representations and perceptual priors.. Eerola and Lahdelma (2021) T. Eerola and I. Lahdelma The anatomy of consonance/dissonance: evaluating acoustic and cultural predictors across multiple datasets with chords. Music & Science 4, p. 1–19. External Links: Document Cited by: Consonance perception and computational models.. Elizalde et al. (2023) B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang CLAP: learning audio concepts from natural language supervision. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, p. 1–5. External Links: Document Cited by: Introduction, Music representations and perceptual priors.. Harrison and Pearce (2020) P. M. C. Harrison and M. T. Pearce Simultaneous consonance in music perception and composition. Psychological Review 127 (2), p. 216–244. External Links: Document Cited by: Introduction, Consonance perception and computational models.. Harte et al. (2006) C. Harte, M. Sandler, and M. Gasser Detecting harmonic change in musical audio. In Proceedings of the 1st ACM Workshop on Audio and Music Computing Multimedia, p. 21–26. External Links: Document Cited by: Introduction, Music representations and perceptual priors.. Huang et al. (2022) Q. Huang, A. Jansen, J. Lee, R. Ganti, J. Y. Li, and D. P. W. Ellis MuLan: a joint embedding of music audio and natural language. In Proceedings of the 23rd International Society for Music Information Retrieval Conference, p. 559–566. Cited by: Introduction, Music representations and perceptual priors.. Hutchinson and Knopoff (1978) W. Hutchinson and L. Knopoff The acoustic component of western consonance. Interface 7 (1), p. 1–29. External Links: Document Cited by: Consonance perception and computational models.. Kaklamani and Simserides (2026) S. Kaklamani and C. Simserides Psychoacoustic study of simple-tone dyads: frequency ratio and pitch. Acoustics 8 (1), p. 14. External Links: Document Cited by: Consonance perception and computational models.. Kang and Herremans (2025) J. Kang and D. Herremans Towards unified music emotion recognition across dimensional and categorical models. External Links: 2502.03979, Document Cited by: Model and data., Music2Emo parameters., Table 14. Korzeniowski and Widmer (2016) F. Korzeniowski and G. Widmer Feature learning for chord recognition: the deep chroma extractor. In Proceedings of the 17th International Society for Music Information Retrieval Conference, p. 37–43. Cited by: Introduction, Music representations and perceptual priors.. Langner (1992) G. Langner Periodicity coding in the auditory system. Hearing Research 60 (2), p. 115–142. External Links: Document Cited by: Introduction, Consonance perception and computational models.. Li et al. (2024) Y. Li, R. Yuan, G. Zhang, Y. Ma, X. Chen, H. Yin, C. Xiao, C. Lin, A. Ragni, E. Benetos, N. Gyenge, R. B. Dannenberg, R. Liu, W. Chen, G. Xia, Y. Shi, W. Huang, Z. Wang, Y. Guo, and J. Fu MERT: acoustic music understanding model with large-scale self-supervised training. In International Conference on Learning Representations, External Links: Link Cited by: Introduction, Introduction, Music representations and perceptual priors., Music representations and perceptual priors., Protocol.. Lin (2004) C. Lin ROUGE: a package for automatic evaluation of summaries. In Text Summarization Branches Out, p. 74–81. Cited by: Protocol.. Liu et al. (2024) S. Liu, A. S. Hussain, C. Sun, and Y. Shan Music understanding LLaMA: advancing text-to-music generation with question answering and captioning. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, p. 286–290. External Links: Document Cited by: Introduction, Music representations and perceptual priors., Protocol., Additional MusicQA Comparisons, Table 13. Lu et al. (2021) W. Lu, J. Wang, M. Won, K. Choi, and X. Song SpecTNT: a time–frequency transformer for music audio. In Proceedings of the 22nd International Society for Music Information Retrieval Conference, p. 396–403. Cited by: Introduction, Music representations and perceptual priors.. Malmberg (1918) C. F. Malmberg The perception of consonance and dissonance. Psychological Monographs 25 (2), p. 93–133. External Links: Document Cited by: Controlled Music-Theory Validation, Intervals. Marjieh et al. (2024) R. Marjieh, P. M. C. Harrison, H. Lee, F. Deligiannaki, and N. Jacoby Timbral effects on consonance disentangle psychoacoustic mechanisms and suggest perceptual origins for musical scales. Nature Communications 15, p. 1482. External Links: Document Cited by: Consonance perception and computational models.. McDermott et al. (2016) J. H. McDermott, A. F. Schultz, E. A. Undurraga, and R. A. Godoy Indifference to dissonance in native amazonians reveals cultural variation in music perception. Nature 535 (7613), p. 547–550. External Links: Document Cited by: Consonance perception and computational models.. Müller and Ewert (2011) M. Müller and S. Ewert Chroma toolbox: matlab implementations for extracting variants of chroma-based audio features. In Proceedings of the 12th International Society for Music Information Retrieval Conference, p. 215–220. Cited by: Music representations and perceptual priors.. Panda et al. (2023) R. Panda, R. Malheiro, and R. P. Paiva Audio features for music emotion recognition: a survey. IEEE Transactions on Affective Computing 14 (1), p. 68–88. External Links: Document Cited by: Recorded music and consonance-aware music AI.. Papineni et al. (2002) K. Papineni, S. Roukos, T. Ward, and W. Zhu BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, p. 311–318. External Links: Document Cited by: Protocol.. Parncutt (1989) R. Parncutt Harmony: a psychoacoustical approach. Springer-Verlag, Berlin. External Links: Document Cited by: Consonance perception and computational models.. Plomp and Levelt (1965) R. Plomp and W. J. M. Levelt Tonal consonance and critical bandwidth. The Journal of the Acoustical Society of America 38 (4), p. 548–560. External Links: Document Cited by: Introduction, Consonance perception and computational models.. Poltronieri et al. (2025) A. Poltronieri, X. Serra, and M. Rocamora From discord to harmony: decomposed consonance-based training for improved audio chord estimation. In Proceedings of the 26th International Society for Music Information Retrieval Conference, Daejeon, South Korea, p. 492–502. External Links: Document Cited by: Introduction, Music representations and perceptual priors.. Schwär et al. (2025) S. Schwär, S. Balke, and M. Müller Measuring sensory dissonance in multi-track music recordings: a case study with wind quartets. In Proceedings of the 26th International Society for Music Information Retrieval Conference, Daejeon, South Korea, p. 117–126. External Links: Document Cited by: Recorded music and consonance-aware music AI.. Schwartz et al. (2003) D. A. Schwartz, C. Q. Howe, and D. Purves The statistical structure of human speech sounds predicts musical universals. The Journal of Neuroscience 23 (18), p. 7160–7168. External Links: Document Cited by: Controlled Music-Theory Validation, Intervals. Sethares (1994) W. A. Sethares Adaptive tunings for musical scales. The Journal of the Acoustical Society of America 96 (1), p. 10–18. Cited by: Amplitude-weighted attribution.. Sethares (2005) W. A. Sethares Tuning, timbre, spectrum, scale. 2 edition, Springer, London. External Links: Document Cited by: Consonance perception and computational models., Amplitude-weighted attribution.. Soleymani et al. (2013) M. Soleymani, M. N. Caro, E. M. Schmidt, C. Sha, and Y. Yang 1000 songs for emotional analysis of music. In Proceedings of the 2nd ACM International Workshop on Crowdsourcing for Multimedia, p. 1–6. External Links: Document Cited by: Model and data.. Stolzenburg (2015) F. Stolzenburg Harmony perception by periodicity detection. Journal of Mathematics and Music 9 (3), p. 215–238. External Links: Document Cited by: Introduction, Introduction, Consonance perception and computational models., Overview, Rational candidates and complexity.. Tenney (1984) J. Tenney John cage and the theory of harmony. In Soundings 13: The Music of James Tenney, P. Garland (Ed.), p. 55–83. Cited by: Introduction, Overview, Rational candidates and complexity.. Touvron et al. (2023) H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama 2: open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. External Links: Document Cited by: Protocol.. von Helmholtz (1954) H. L. F. von Helmholtz On the sensations of tone as a physiological basis for the theory of music. Dover Publications, New York. Note: Second English edition, translated by Alexander J. Ellis; original German work published in 1863 Cited by: Introduction, Consonance perception and computational models.. Wagner et al. (2022) B. Wagner, S. Czoschke, B. Tillmann, S. Koelsch, P. Vuust, and E. Brattico Pitch chroma information is processed in addition to pitch height information with more than two pitch-range categories. Attention, Perception, & Psychophysics 84 (5), p. 1757–1771. External Links: Document Cited by: Continuous pitch intervals and octave folding.. Wand (2012) A. Wand On the conception and measure of consonance. Leonardo Music Journal 22, p. 73–78. External Links: Document Cited by: Consonance perception and computational models.. Weck et al. (2026) B. Weck, P. Puentes, A. Poltronieri, S. Prabhu, and D. Bogdanov HumMusQA: a human-written music understanding QA benchmark dataset. In Proceedings of the 4th Workshop on NLP for Music and Audio (NLP4MusA 2026), Rabat, Morocco, p. 58–67. External Links: Document Cited by: Music representations and perceptual priors.. Won et al. (2024) M. Won, Y. Hung, and D. Le A foundation model for music informatics. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, p. 1226–1230. External Links: Document Cited by: Introduction, Music representations and perceptual priors.. Xia et al. (2026) H. Xia, Z. Huang, Y. Tan, and S. Song Let the model learn to feel: mode-guided tonality injection for symbolic music emotion recognition. Proceedings of the AAAI Conference on Artificial Intelligence 40 (3), p. 2182–2190. External Links: Document Cited by: Introduction. Zhang et al. (2018) K. Zhang, H. Zhang, S. Li, C. Yang, and L. Sun The PMEmo dataset for music emotion recognition. In Proceedings of the 2018 ACM International Conference on Multimedia Retrieval, p. 135–142. External Links: Document Cited by: Model and data.. Zhang et al. (2020) T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi BERTScore: evaluating text generation with BERT. In International Conference on Learning Representations, Cited by: Protocol..