Paper deep dive
Assessing AI-generated music detection in real-world broadcast monitoring
David LĂłpez-Ayala, Fernando GarcĂa de la Cruz, Pablo Zinemanas, Emilio Molina, MartĂn Rocamora
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The proliferation of AI-generated music in broadcast media raises concerns about transparency and fair compensation, but reliable detection under real broadcast conditions remains unresolved. Existing studies report substantial performance degradation in this domain, yet their evaluations are limited to synthetic broadcast data. To address this gap, we introduce BAMM (Broadcast AI-Music Monitoring), a 40-hour dataset of real-world television recordings containing AI-generated and human-made music. We compare clean-trained and broadcast-trained CNN variants across three progressively more challenging scenarios: Clean Foreground Music (CFM), Synthetic TV Broadcast (STB), and Real TV Broadcast (RTB). Both models achieve near-perfect performance on CFM but degrade substantially under synthetic broadcast conditions. Broadcast-oriented training improves robustness compared with clean training, although performance remains limited. On RTB, evaluated using BAMM, both models degrade further and show substantial score overlap between AI-generated and human-made music. These results expose a critical domain gap and show that current training approaches on CNN-based detectors remain insufficient for reliable AI-generated music detection in broadcast monitoring.
Tags
Links
- Source: https://arxiv.org/abs/2608.07359v1
- Canonical: https://arxiv.org/abs/2608.07359v1
Trouble viewing inline? Open PDF directly â
Full Text
33,183 characters extracted from source content.
Expand or collapse full text
ASSESSING AI-GENERATED MUSIC DETECTION IN REAL-WORLD BROADCAST MONITORING David LĂłpez-Ayala 1 Fernando Garcia de la Cruz 1 Pablo Zinemanas 2 Emilio Molina 2 MartĂn Rocamora 1 1 Music Technology Group, Universitat Pompeu Fabra, Barcelona, Spain 2 BMAT Licensing S.L., Barcelona, Spain jorgedavid.lopez@upf.edu ABSTRACT The proliferation of AI-generated music in broadcast me- dia raises concerns about transparency and fair compen- sation, but reliable detection under real broadcast con- ditions remains unresolved. Existing studies report sub- stantial performance degradation in this domain, yet their evaluations are limited to synthetic broadcast data. To address this gap, we introduce BAMM (Broadcast AI- Music Monitoring), a 40-hour dataset of real-world televi- sion recordings containing AI-generated and human-made music. We compare clean-trained and broadcast-trained CNN variants across three progressively more challeng- ing scenarios: Clean Foreground Music (CFM), Synthetic TV Broadcast (STB), and Real TV Broadcast (RTB). Both models achieve near-perfect performance on CFM but de- grade substantially under synthetic broadcast conditions. Broadcast-oriented training improves robustness compared with clean training, although performance remains limited. On RTB, evaluated using BAMM, both models degrade further and show substantial score overlap between AI- generated and human-made music. These results expose a critical domain gap and show that current training ap- proaches on CNN-based detectors remain insufficient for reliable AI-generated music detection in broadcast moni- toring. 1. INTRODUCTION Generative music models have rapidly lowered the bar- rier to music production. Their widespread adoption has brought to light a range of technical and legal challenges concerning training data provenance, training data replica- tion, the definition of authorship, and compensation frame- works. These tensions have resulted in lawsuits against AI-music companies such as Suno and Udio [1], as well as licensing deals between generative-music platforms and rightsholders [2]. In parallel, policies such as the EU AI © D. LĂłpez-Ayala, F. Garcia de la Cruz, P Zinemanas, E Molina, and M. Rocamora. Licensed under a Creative Commons Attri- bution 4.0 International License (C BY 4.0). Attribution: D. LĂłpez- Ayala, F. Garcia de la Cruz, P Zinemanas, E Molina, and M. Rocamora, âAssessing AI-generated music detection in real-world broadcast moni- toringâ, in Proc. of the 27th Int. Society for Music Information Retrieval Conf., Abu Dhabi, UAE, 2026. Figure 1. Generation pipeline for BAMM dataset. Act introduce transparency requirements for AI-generated content to make AI-generated media identifiable [3]. The need for robust detection tools has become in- creasingly urgent as synthetic content reaches an indus- trial scale. This urgency is underscored by Deezerâs report that, for the first time, 50% of its daily uploads were de- tected as fully AI-generated. [4]. Similar trends have also been observed in broadcast media, with companies such as BMAT reporting a growing presence of AI-generated mu- sic in television recordings. [5]. While existing AI-generated music detectors achieve near-perfect accuracy on isolated, high-fidelity tracks, they are often not robust to domain shifts and common au- dio transformations, such as random pitch shifting, time stretching, EQ, reverb, among others [6].The broad- cast environment introduces significant acoustic complex- ities: music segments are typically short, frequently back- grounded by dominant speech or sound effects, and sub- jected to transmission constraints. Recent benchmarks have demonstrated that, even in synthetic mixtures, these factors substantially degrade performance on state-of-the- arXiv:2608.07359v1 [eess.AS] 7 Aug 2026 art architectures [7]. However, the extent to which AI de- tectors work in real-world broadcast conditions remains under-explored. To bridge this gap, we introduce BAMM, a 40-hour dataset of real TV broadcast content specifically curated for AI-generated music detection. The dataset is balanced, containing approximately 20 hours of AI-generated mu- sic and 20 hours of human-made music, and was built by matching clean references to their broadcast occurrences through audio fingerprinting. The dataset contains clips between 5 s and 60 s, with music appearing either in the foreground or background. By collecting material from real TV content aired worldwide, BAMM captures the variability and practical constraints of broadcast environ- ments. We leverage this dataset to evaluate state-of-the- art CNN models, and identify current architectural limita- tions, and thus provide a reference evaluation setting for real-world AI music identification. The .mp4 recordings are publicly available on Zenodo, and the baseline code and evaluation scripts are in Github 1 . 2. RELATED WORK The detection of AI-generated music has emerged within the MIR field as part of a broader effort to identify syn- thetic audio, alongside related tasks such as voice spoofing and synthetic speech detection [8, 9]. Music-specific work has primarily focused on unintentionally left artifacts by generative systems. Afchar et al. [6] showed that these arti- facts introduced during the decoding stage can be detected, under controlled conditions, with CNN-based models, us- ing neural audio codecs such as Encodec and DAC [10, 11] to simulate this process. In later work, Afchar et al. [12] further characterized these artifacts through Fourier theory and showed that similar patterns are also present in mu- sic generated by commercial platforms such as Suno and Udio. Then Rahman et al. [13] introduced SONICS, a dataset composed of songs generated with Suno and Udio for the AI-generated class and samples from the Genius Lyrics Dataset [14] for the human-made class. They also pro- posed the SpecTTTra family of models, demonstrating that spatio-temporal representations are effective for distin- guishing AI-generated from human-made music, achiev- ing near-perfect performance. However, despite promis- ing results, existing approaches remain largely limited to controlled environments. Cros Vila et al. [15] showed that variations in sampling rate and bit rate may encour- age detectors to rely on dataset-specific shortcuts rather than features intrinsic to AI-generated music, while mod- els with near-perfect in-domain performance may general- ize poorly beyond their training conditions. These works establish strong baselines, constraints, and limitations for the task, showcasing that robust AI- generated music detection remains unresolved in more complex scenarios. One such scenario is broadcast, which introduces substantially different conditions: music may 1 https://github.com/DaveLoay/BAMM-dataset.git appear in the foreground or background and may be mixed with speech and sound effects at different relative loud- ness levels. Previous work has addressed these challenges through distinct tasks. Relative music loudness estimation aims to characterize how music appears within a broadcast mix [16], with OpenBMAT providing a dataset specifically designed for this purpose [17]. Audio fingerprinting, in contrast, focuses on matching reference tracks against TV recordings [18], with BAF providing a benchmark for eval- uating this task under broadcast conditions [19]. Together, these works highlight the prevalence of background music and mixed-audio conditions in real broadcast content. LĂłpez-Ayala et al. [7] addressed AI-generated music detection in broadcast settings using synthetically con- structed television content. Their study showed that ex- isting detectors experience substantial performance degra- dation when music appears in short excerpts, is masked by dominant speech or sound effects, or is subject to broadcast-specific production and storage constraints. These include transitions, editing practices, changes in dy- namics, and low-quality audio encoding, as further dis- cussed in Section 4.2. To support this evaluation, they introduced AI-OpenBMAT, a dataset designed for AI- generated music detection in broadcast settings. It repro- duces the duration patterns and loudness relationships of real television audio by combining human-made produc- tion music with stylistically matched continuations gen- erated using Suno. Although AI-OpenBMAT provides a valuable controlled benchmark, its broadcast conditions are still simulated through synthetic mixtures. This work extends that line of research by introducing a dataset built from real television broadcast recordings and using it to evaluate current AI-generated music detectors under prac- tical broadcast conditions. 3. BAMM DATASET BAMM (Broadcast AI-Music Monitoring) is a novel dataset comprising 40 hours of real-world broadcast con- tent curated for AI-generated music detection. The dataset was constructed by identifying clean AI-generated and human-made reference tracks within a global broadcast archive using audio fingerprinting, followed by a multi- stage filtering process to retain valid broadcast occur- rences. BAMM consists of approximately 20 hours of AI-generated music and 20 hours of human-made music, stored as .mp4 clips. Each sample captures original broad- cast audio and video, preserving the authentic acoustic degradations and production contexts (e.g., music in the background, sound effects) found in practice. The clips were sampled from a global broadcast archive that mon- itors over 4,200 television channels worldwide, with all samples aired between January 2025 and March 2026 and durations ranging from 5 s to 60 s. Reflecting the tech- nical constraints of large-scale industrial monitoring, the audio is monophonic, sampled at 8 kHz, and encoded with AAC-LC at bitrates of at least 40 kbps. While lower than consumer broadcast standards, these specifications are rep- resentative of the low-bitrate proxy streams used in global industrial monitoring, posing a significant worst-case chal- lenge for detection models. To ensure trustworthy ground truth labels, we developed an automated pipeline to bridge the gap between detecting AI content in clean conditions and real-world broadcast. We first curated a reference corpus of clean, foreground .mp3 tracks for both the AI-generated and human-made classes. We then leveraged an audio fingerprinting system to identify the precise occurrences of these tracks within the global broadcast archive. The retrieved segments were filtered using a multi-stage process devised to isolate valid clips that satisfied the technical requirements described above, as illustrated in Fig. 1. The following subsections describe the individual stages of this procedure. 3.1 Reference Track Selection The construction pipeline begins with a candidate pool of clean MP3 reference tracks. Before searching for occur- rences of these tracks in TV broadcast content, we cate- gorize them using an ensemble of five AI-music detectors, hereafter referred to as the detector ensemble. This config- uration promotes methodological diversity by combining detectors trained on different datasets, operating at differ- ent sampling rates, and relying on distinct feature repre- sentations. The labels were determined separately for both classes. For the human-made class, reference tracks were restricted to releases between January 2020 and December 2022, ensuring that they predated Suno v3.5 and other publicly available or commercial generative music systems capa- ble of producing music of comparable quality. For the AI- generated class, a candidate track was included only if it exceeded the calibrated decision thresholds across all five detectors, requiring unanimous agreement from the ensem- ble. Regarding the models in the detector ensemble. The first model is a publicly available convolutional neural net- work presented and assessed in this paper under the name CNN Clean. It is based on the architecture proposed by Afchar et al. [6], was trained on publicly available data, It operates on 5-second audio windows represented as mel- spectrograms computed from audio resampled to 8 kHz. The model weights and inference code are available in the project repository. For more in-depth implementation de- tails, refer to Section 4.1. The second model is the publicly available SpecTTTra alpha-5s detector [13]. It was selected because it achieved the best performance among the evaluated SpecTTTra vari- ants on our control dataset. The model operates on spatio- temporal tokens extracted from audio resampled to 16 kHz and was trained on the SONICS suno v3.5 subset. The remaining three models of the detector ensemble were developed for internal research purposes and trained on in-house data. The third model is a CNN based on the architecture proposed by Afchar et al. [6]. It operates on 5-second audio windows represented as mel-spectrograms computed from audio resampled to 16 kHz. The fourth model is a logistic regression classifier following the ap- 0.00.20.40.60.81.0 AI Probability Score H1 2021 H2 H1 2022 H2 H1 2023 H2 H1 2024 H2 H1 2025 H2 H1 2026 Suno v3.5 (Summer 2024) Suno v4 (Nov 2024) Suno v4.5 (May 2025) Suno v4.5+ (Jul 2025) Suno v5 (Sep 2025) Figure 2. Distribution of AI-probability scores produced by the detector ensemble for tracks released from 2021 to 2026. Score values remain low before the release of Suno v3.5, and increase afterward, indicating the appearance of Suno-generated music. proach introduced by Afchar et al. [12]. It operates on artifact-fingerprint features extracted from audio resam- pled to 16 kHz. The fifth model is an MLP trained on the same feature representation as the logistic regression classifier. All five models were trained as binary classifiers under clean foreground music conditions to distinguish human- made from AI-generated tracks (Suno v3.5), and all achieved F1 scores above 98%. We performed a conser- vative calibration on each model to achieve a zero false positive rate. The calibration set included 5,000 tracks from Da-TACOS [20], providing a verified human base- line that predates modern generative models, and a control set of 500 AI-generated tracks from the clean subset of AI- OpenBMAT. By deriving model-specific decision thresh- olds that yielded zero errors on the human-composed set, we obtained a labeling criterion with high confidence in the resulting AI-generated samples. To further validate the selection strategy, we analyzed the temporal evolution of AI-positive detections among the artists with the highest numbers of positive cases, as shown in Figure 2. Detection scores are consistently low before the Summer 2024 release date of Suno v3.5, in- crease markedly after its introduction, and later decrease over time. This behavior is consistent with the real-world adoption of Suno v3.5 and the subsequent appearance of newer generative models from the same family. 3.2 Fingerprint & Broadcast Retrieval Once the reference tracks for AI-generated and human- made classes were defined, we used a private implemen- tation of audio fingerprinting based on landmarks [21] to locate matches in our TV broadcast archive. The matched segments were searched in recordings aired between Jan- uary 2025 and March 2026 and then extracted directly from the original TV emissions. These recordings are archived at a sampling rate of 8 kHz and a bitrate of 40 kbps. 3.3 DMD Filtering Following fingerprint retrieval, we applied a Deep Mu- sic Detector (DMD), based on the approach proposed by MelĂ©ndez et al. [16]. This model classifies audio segments as foreground music, background music, or speech. We kept only clips labeled as containing music, whether in the foreground or background, thereby removing sound effects and incorrect fingerprint matches. These foreground and background labels are also used when analyzing the per- formance of AI-generated music detection models (Sec- tion 5.3). 4. EXPERIMENTAL SETUP This benchmark examines how the training domain affects CNN-based detectors designed to identify AI-generated music in real broadcast settings. We compare two vari- ants of the architecture proposed by Afchar et al. [6], both operating on 5-second windows but trained under different conditions: clean foreground music and broadcast-oriented data. This comparison isolates the effect of the training domain while keeping the model capacity fixed, allowing us to examine the limitations of CNN-based AI-generated music detectors in broadcast environments. The bench- mark focuses on Suno v3.5, as model-agnostic detection remains an open problem, and this generative model has publicly available data and matched reference detectors. 4.1 Shared Architecture and Preprocessing The CNN-based model has six convolutional layers with filter sizes [16, 32, 64, 128, 256, 512] and kernel size 3, fol- lowed by average pooling and two fully connected layers. In the pre-processing pipeline, audio is converted to mono and resampled to 8 kHz. Then the waveform is transformed into a time-frequency representation using an STFT with a window size of 2048 and a hop length of 1024. The resulting power spectrogram is projected onto 128 Mel bands spanning 20 Hz to 4 kHz, converted to the logarithmic dB scale, and standardized with a fixed global mean ofâ4.0 and standard deviation of 3.0. The Mel- spectrogram is then segmented into discrete time windows to form the input tensor of the CNN. 4.1.1 CNN Clean CNN Clean is trained on clean foreground music only. Au- dio is converted to mono and resampled to 8 kHz in the pre-processing pipeline. The human-made class is drawn from the FMA-medium dataset, comprising approximately 25k songs, while the AI-generated class is built from the Suno v3.5 subset of SONICS, comprising approximately 19k songs. For each batch, five 5 s snippets are randomly sampled from each track, enabling consistent training. 4.1.2 CNN Broadcast CNN Broadcast uses the same architecture and input repre- sentation as CNN Clean, but is trained on data designed to emulate broadcast conditions by mixing music and speech. The music sources are the same as for CNN Clean, namely FMA-medium for the human-made class and the Suno v3.5 subset of SONICS for the AI-generated class. Speech is taken from the 360-hour clean training partition of Lib- riSpeech [22]. The training set is constructed with a controlled ratio of 70% mixed samples and 30% clean samples for each class. In the mixed condition, speech segments are concatenated to span the full duration of each music track and mixed with the music signal at a randomly sampled SNR between â30 dB and +30 dB. The resulting files are exported in mono at 8 kHz and encoded as .mp3 at 40 kbps before entering the shared preprocessing pipeline. 4.2 Evaluation Scenarios To assess the benchmark models under progressively more challenging conditions, we evaluate them across three sce- narios: (i) Clean Foreground Music (CFM), (i) Synthetic TV Broadcast (STB), and (i) Real TV Broadcast (RTB). 4.2.1 Clean Foreground Music (CFM). For this scenario, we evaluate the models on the clean mu- sic tracks used to build AI-OpenBMAT [7]. The dataset contains 476 humanâAI pairs with a sampling rate of 22.05 kHz and a bitrate of 353 kbps. The AI tracks were generated from their corresponding human tracks using Sunoâs extend function. As a result, the pairs are closely matched in musical characteristics such as style, timbre, pitch, and key. This scenario provides a direct point of comparison with the synthetic broadcast scenario, which is built from the same source material. 4.2.2 Synthetic TV Broadcast (STB). We evaluate the synthetic broadcast scenario using the full AI-OpenBMAT dataset, following [7]. The dataset con- tains 3,294 tracks, totaling 54.9 hours of synthetic broad- cast material, with a sampling rate of 22.05 kHz and a bi- trate of 353 kbps. These data were created from the same clean foreground references described above, using con- trolled mixtures designed to accurately emulate broadcast conditions. This scenario allows us to evaluate model per- formance under synthetic broadcast conditions while keep- ing the underlying musical material controlled. 4.2.3 Real TV Broadcast (RTB). The real broadcast scenario uses the full BAMM dataset introduced in this work. It comprises 40 hours of real TV broadcast recordings, with approximately 20 hours per class, at a sampling rate of 8 kHz and a bitrate of 40 kbps using AAC-LC encoding. This setting extends previous broadcast-oriented evaluations by shifting from synthetic ScenarioCNN CleanCNN Broadcast F1ROCF1ROC CFM0.9920.9980.9920.996 STB0.3420.9090.6610.926 RTB0.1860.7070.4720.775 Table 1. Results across the three evaluation scenarios: Clean Foreground Music (CFM), Synthetic TV Broadcast (STB), and Real TV Broadcast (RTB). approximations to real television content, enabling the models to be tested under the conditions of the target envi- ronment. 5. RESULTS We evaluate both models across three scenarios with in- creasing levels of broadcast complexity: Clean Fore- ground Music (CFM), Synthetic TV Broadcast (STB), and Real TV Broadcast (RTB). The comparison between CNN Clean and CNN Broadcast allows us to analyze how the training domain influences the robustness of AI-generated music detection. While the CFM scenario provides a con- trolled evaluation of artifact detection, the STB setup in- troduces synthetic broadcast conditions through controlled musicâspeech mixtures. Finally, the RTB case evaluates both models on the BAMM dataset introduced in this work, capturing the challenges of real television content. 5.1 Evaluation on Clean Foreground Music The first stage of this evaluation aims to verify whether de- tectors trained under real-world broadcast constraints, i.e., low-quality audio at 8 kHz sampling rate, remain reliable for detecting AI-generated music under clean, high-quality foreground conditions. As shown in Table 1, both CNN Clean and CNN Broadcast achieve F1-scores above 99%, demonstrating that the artifacts learned from low-quality training data are still present in the high-quality data typ- ically used for this task and remain detectable even when considering only the lower-frequency components of the audio spectrum. These results validate the proposed mod- els and provide a baseline for the subsequent evaluations. 5.2 Evaluation on Synthetic TV Broadcast The synthetic TV broadcast (STB) scenario represents the first evaluation under mixed-audio conditions, in which AI-generated music is no longer isolated but combined with speech. As observed in Table 1, the transition from clean foreground music to synthetic broadcast conditions produces a significant reduction in detection performance for both models.The clean-trained model suffers the greatest degradation, demonstrating that artifacts learned from isolated music do not fully transfer when partially masked by the surrounding broadcast context. This re- sult is consistent with the performance reported for AI- OpenBMAT [7] and validates the experimental setup. The CNN Broadcast model alleviates this degradation by learn- ing from similar mixed conditions, demonstrating that ex- posure to broadcast-like data improves artifact detection under masking. However, its limited performance suggests that the problem is not only related to the training distribu- tion but also to an intrinsic reduction in artifact visibility arising from the broadcast scenarioâs mixture process. 5.3 Evaluation on Real TV Broadcast The RTB scenario is the most challenging setting in this work, as it evaluates the detectors directly on real television content. Unlike STB, which approximates broadcast con- ditions using controlled synthetic mixtures, RTB includes the full complexity of real recordings. As shown in Ta- ble 1, both models perform worse than in STB, indicating that the synthetic setting does not fully capture real broad- cast conditions. The CNN Clean model shows the largest degradation, confirming that artifacts learned from isolated foreground music do not transfer reliably to real broadcasts. In ad- dition to the reduced audio quality of the RTB recordings (8 kHz sampling rate and 40 kbps AAC-LC), sound ef- fects, transitions, and diverse mixing strategies can further obscure AI-related artifacts. Figure 3 shows a detailed comparison of the modelsâ performance when considering BAMM clips with music in the background or in the foreground. Both models re- tain some discriminative ability and perform better on fore- ground than on background music. This confirms that masking is a major source of degradation. However, the limited performance on foreground segments also indicates that degraded audio quality and other broadcast sounds af- fect artifact visibility. The score distributions in Figure 4 show that the main difficulty is detecting AI-generated music rather than re- jecting human content. While human samples are gener- ally assigned scores near zero, many AI samples also re- ceive low scores, particularly in background music. Over- all, broadcast-simulated training improves robustness, but it does not fully overcome the effects of real broadcast con- ditions. 6. CONCLUSIONS In this paper, we introduced BAMM, a dataset for AI- generated music detection designed to address key limi- tations of existing benchmarks. Unlike datasets focused primarily on clean or isolated musical excerpts, BAMM targets the real broadcast setting, where music appears in short segments, alternates between foreground and back- ground roles, and is mixed with speech and sound effects. Our experiments show that CNN-based detectors trained on clean foreground music or simulated broadcast data retain some discriminative capacity under controlled conditions, but their performance degrades substantially as the evaluation scenario becomes more realistic. In partic- ular, although performance remains limited for both fore- ground and background music, the models perform slightly 0.00.20.40.60.81.0 False Positive Rate 0.0 0.2 0.4 0.6 0.8 1.0 True Positive Rate CNN Broadcast (AUC=0.775) CNN Clean (AUC=0.707) Random 0.00.20.40.60.81.0 False Positive Rate 0.0 0.2 0.4 0.6 0.8 1.0 CNN Broadcast (AUC=0.858) CNN Clean (AUC=0.782) Random 0.00.20.40.60.81.0 False Positive Rate 0.0 0.2 0.4 0.6 0.8 1.0 Background music CNN Broadcast (AUC=0.745) CNN Clean (AUC=0.667) Random All ClipsForeground music Figure 3. ROC curves for the clean-trained and broadcast-trained CNN models when evaluated on all BAMM clips, only on foreground music, and only on background music. Performance decreases for background music, showing that detection becomes more difficult when the musical signal is less salient or mixed with other content. 0 200 400 600 800 1000 1200 1400 Music CNN Clean Human AI 0 200 400 600 800 1000 1200 CNN Broadcast Human AI 0.00.20.40.60.81.0 Probability (AI) 0 500 1000 1500 2000 2500 Background music CNN Clean Human AI 0.00.20.40.60.81.0 Probability (AI) 0 250 500 750 1000 1250 1500 1750 2000 CNN Broadcast Human AI Figure 4. Score distributions for the CNN models under Real TV Broadcast conditions. Blue and red densities de- note human-made and AI-generated clips. better when the music is in the foreground. This sug- gests that current detectors are not reliably capturing AI- generated music artifacts, and that the task becomes even more difficult when the musical signal is less salient or em- bedded within complex audio mixtures. These findings highlight the need for more robust detec- tion methods capable of modeling the acoustic and contex- tual complexity of real broadcast media. BAMM provides a benchmark for this direction by exposing failure modes not fully captured by clean-music datasets and encouraging the development of models that can detect AI-generated music in challenging, real-world broadcast environments. 7. ACKNOWLEDGMENTS This work is supported by the "CĂĄtedra IA y MĂșsica" project (TSI-100929-2023-1),funded by the Sec- retarĂa de Estado de DigitalizaciĂłn e Inteligencia Artificial, the European Union-Next Generation EU funds and BMAT Music Innovators.And by the "IMPA" project (PID2023-152250OB-I00) funded by MCIU/AEI/10.13039/501100011033/FEDER, UE. 8. REFERENCES [1] D. Tencer, âMajor record companies sue Suno, Udioforâmassinfringementâofcopyright,â https://w.musicbusinessworldwide.com/major- record-companies-sue-ai-music-generators-suno-udio- for-mass-infringement-of-copyright/, Jun. 2024. [2] M.Stassen,âWarnerMusicGroupstrikes âlandmarkâdealwithSuno;settlescopy- rightlawsuitagainstAImusicgenerator,â https://w.musicbusinessworldwide.com/warner- music-group-settles-with-suno-strikes-first-of-its- kind-deal-with-ai-song-generator/, Nov. 2025. [3] X. Serra, R. O. Araz, R. Batlle-Roca, L. Juvela, D. LĂłpez, and M. Rocamora, âTechnical Solutions for Marking and Detecting AI-generated Audio in the Context of Article 50 of the AI Act,â Euro- pean Commission, Technical Report, 2026, tender EC- CNECT/2024/VLVP/0115 Watermarking Audio. [4] J. Wendel, âAi music tops 50% of daily uploads on deezer,âhttps://newsroom-deezer.com/2026/07/ ai-music-exceeds-50-percent-daily-uploads-deezer/, 2026, [Online; accessed 24-July-2026]. [5] D. Bolboac Ì a, âFirst, video killed the radio star, and now AI is going after video: A study on detect- ing GenAI music in broadcast audio,â https://w. bmat.com/genai-music-broadcasting/, 2026, [Online; accessed 24-July-2026]. [6] D. Afchar, G. Meseguer-Brocal, and R. Hennequin, âAI-generated music detection and its challenges,â in Proc. of the IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP), 2025. [7] D. LĂłpez-Ayala, A. Cabello, P. Zinemanas, E. Molina, and M. Rocamora, âAI-generated music detection in broadcast monitoring,â in Proc. of the 2026 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026. [8] Y. Li, M. Milling, L. Specia, and B. W. Schuller, âFrom Audio Deepfake Detection to AI-Generated Music De- tection â a pathway and overview,â arXiv preprint arXiv:2412.00571, 2024. [9] D. Salvi, H. V. Koops, and E. Quinton, âNot All Deep- fakes Are Created Equal: Triaging Audio Forgeries for Robust Deepfake Singer Identification,â arXiv preprint arXiv:2510.17474, 2025. [10] A. DĂ©fossez, J. Copet, G. Synnaeve, and Y. Adi, âHigh fidelity neural audio compression,â Transactions on Machine Learning Research (TMLR), 2023. [11] R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, âHigh-fidelity audio compression with im- proved RVQGAN,â in Proc. of the Conference on Neu- ral Information Processing Systems (NeurIPS), 2023. [12] D. Afchar, G. Meseguer-Brocal, K. Akesbi, and R. Hennequin, âA Fourier Explanation of AI-music Artifacts,â in Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2025. [13] M. A. Rahman, Z. I. A. Hakim, N. H. Sarker, B. Paul, and S. A. Fattah, âSonics: Synthetic or not - identify- ing counterfeit songs,â in International Conference on Learning Representations (ICLR), 2025. [14] C.GDCJ,âGeniussonglyrics,âhttps: //w.kaggle.com/datasets/carlosgdcj/ genius-song-lyrics-with-language-information, 2023, [Online; accessed 16-March-2026]. [15] L. C. Vila, B. L. T. Sturm, L. Casini, and D. Dal- mazzo, âThe AI Music Arms Race: On the Detec- tion of AI-Generated Music,â Transactions of the In- ternational Society for Music Information Retrieval (TISMIR), vol. 8, no. 1, p. 179â194, 2025. [16] B. MelĂ©ndez CatalĂĄn, âRelative music loudness esti- mation in TV broadcast audio using deep learning: An industrial perspective,â Ph.D. dissertation, Universitat Pompeu Fabra, Apr. 2021. [17] B. MelĂ©ndez-CatalĂĄn, E. Molina, and E. GĂłmez, âOpen broadcast media audio from tv: A dataset of tv broadcast audio with relative music loudness anno- tations,â Transactions of the International Society for Music Information Retrieval (TISMIR), vol. 2, no. 1, p. 43â51, 2019. [18] G. CortĂšs-SebastiĂ , M. Miron, E. Molina, A. Ciurana, and X. Serra, âEnhanced television broadcast monitor- ing with source separation-assisted audio fingerprint- ing: A case study,â Multimedia Tools and Applications, vol. 84, no. 42, p. 50 595â50 628, 2025. [19] G. CortĂšs, A. Ciurana, E. Molina, M. Miron, O. Mey- ers, J. Six, and X. Serra, âBAF: An audio fingerprint- ing dataset for broadcast monitoring,â in Proc. of the International Society for Music Information Retrieval Conference (ISMIR), 2022. [20] F. Yesiler, C. Tralie, A. Correya, D. F. Silva, P. Tovsto- gan, E. GĂłmez, and X. Serra, âDa-TACOS: A dataset for cover song identification and understanding,â in Proc. of the International Society for Music Informa- tion Retrieval Conference (ISMIR), 2019. [21] G. CortĂšs SebastiĂ , âMusic identification with audio fingerprinting an industrial perspective,â Ph.D. disser- tation, Universitat Pompeu Fabra, Feb. 2025. [22] M. Korvas, O. PlĂĄtek, O. DuĆĄek, L. Ćœilka, and F. Ju- r Ë cĂ Ë cek, âFree English and Czech telephone speech cor- pus shared under the C-BY-SA 3.0 license,â in Pro- ceedings of the International Conference on Language Resources and Evaluation (LREC), 2014.