Paper deep dive
YMIR: A new Benchmark Dataset and Model for Arabic Yemeni Music Genre Classification Using Convolutional Neural Networks
Moeen AL-Makhlafi, Abdulrahman A. AlKannad, Eiad Almekhlafi, Nawaf Q. Othman Ahmed Mohammed, Saher Qaid
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 99%
Last extracted: 4/10/2026, 2:54:49 AM
Summary
The paper introduces the Yemeni Music Information Retrieval (YMIR) dataset, a collection of 1,475 expert-annotated audio clips across five traditional Yemeni genres (Sanaani, Hadhrami, Lahji, Tihami, and Adeni). It also proposes the Yemeni Music Classification Model (YMCM), a CNN-based architecture that achieves 98.8% accuracy using Mel-spectrogram features, outperforming standard models like AlexNet, VGG16, and MobileNet.
Entities (5)
Relation Signals (3)
YMCM â classifies â YMIR
confidence 100% · YMCM is a system designed to classify music genres from the YMIR dataset.
YMIR â containsgenre â Sanaani
confidence 100% · YMIR dataset contains 1,475 audio clips covering five traditional Yemeni genres: Sanaani...
YMCM â usesfeature â Mel-spectrogram
confidence 100% · YMCM achieves the highest accuracy of 98.8% with Mel-spectrogram features.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Automatic music genre classification is a major task in music information retrieval; however, most current benchmarks and models have been developed primarily for Western music, leaving culturally specific traditions underrepresented. In this paper, we introduce the Yemeni Music Information Retrieval (YMIR) dataset, which contains 1,475 carefully selected audio clips covering five traditional Yemeni genres: Sanaani, Hadhrami, Lahji, Tihami, and Adeni. The dataset was labeled by five Yemeni music experts following a clear and structured protocol, resulting in strong inter-annotator agreement (Fleiss kappa = 0.85). We also propose the Yemeni Music Classification Model (YMCM), a convolutional neural network (CNN)-based system designed to classify music genres from time-frequency features. Using a consistent preprocessing pipeline, we perform a systematic comparison across six experimental groups and five different architectures, resulting in a total of 30 experiments. Specifically, we evaluate several feature representations, including Mel-spectrograms, Chroma, FilterBank, and MFCCs with 13, 20, and 40 coefficients, and benchmark YMCM against standard models (AlexNet, VGG16, MobileNet, and a baseline CNN) under the same experimental conditions. The experimental findings reveal that YMCM is the most effective, achieving the highest accuracy of 98.8% with Mel-spectrogram features. The results also provide practical insights into the relationship between feature representation and model capacity. The findings establish YMIR as a useful benchmark and YMCM as a strong baseline for classifying Yemeni music genres.
Tags
Links
- Source: https://arxiv.org/abs/2604.05011v1
- Canonical: https://arxiv.org/abs/2604.05011v1
Trouble viewing inline? Open PDF directly â
Full Text
47,203 characters extracted from source content.
Expand or collapse full text
JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20211 YMIR: A new Benchmark Dataset and Model for Arabic Yemeni Music Genre Classification Using Convolutional Neural Networks Moeen AL-Makhlafi, Eiad Almekhlafi, Abdulrahman A. AlKannad, Nawaf Q. Othman, Ahmed Mohammed, and Saher Qaid AbstractâAutomatic music genre classification is a major task in music information retrieval; however, most current bench- marks and models have been developed primarily for Western music, leaving culturally specific traditions underrepresented. In this paper, we introduce the Yemeni Music Information Retrieval (YMIR) dataset, which contains 1,475 carefully selected audio clips covering five traditional Yemeni genres: Sanaâani, Hadhrami, Lahji, Tihami, and Adeni. The dataset was labeled by five Yemeni music experts following a clear and structured protocol, resulting in strong inter-annotator agreement (Fleissâ Îș= 0.85). We also propose the Yemeni Music Classification Model (YMCM), a convolutional neural network (CNN)-based system designed to classify music genres from time-frequency features. Using a consistent preprocessing pipeline, we perform a systematic comparison across six experimental groups and five different architectures, resulting in a total of 30 experiments. Specifically, we evaluate several feature representations, Mel- spectrograms, Chroma, FilterBank, and MFCCs with 13, 20, and 40 coefficients, and benchmark YMCM against standard models (AlexNet, VGG16, MobileNet, and a baseline CNN) under the same experimental conditions. The experimental findings reveal that YMCM is the most effective, achieving the highest accuracy of 98.8 % with Mel-spectrogram features. The experimental findings not only reveal that YMCM is the most effective but also provide practical insights into the relationship between feature representation and model capacity. The findings put YMIR at a good benchmark and YMCM at a strong baseline for classifying the genre of Yemeni music. Index TermsâMusic genre classification, Audio classification, Convolutional neural networks (CNNs), Benchmark dataset, Mel- spectrogram. I. INTRODUCTION Music is culturally and socially significant, taking on diverse forms and styles around the world. The wide range of personal music tastes makes precise music classification a core task for creating personalized recommendations [1]. Since song titles alone are not enough tools for classification, musical genre has proven to be one of the most reliable criteria for this task [2]. Although genre is often assigned by hand, it remains essential for matching recommendations to listenersâ tastes. Major services like Spotify and SoundCloud use genre to classify music, helping them engage users with personalized content [3]. The exponential growth of multimedia content across a wide range of digital platforms has substantially exacerbated the difficulties involved in effectively indexing, browsing, M. AL-Makhlafi, A. AlKannad, N. Othman A. Mohammed and S. Qaid are School of Artificial Intelligence, Xidian University, Xiâan, China. E. Almekhlafi is with Department of Information Science and Technology, Northwest University, Xiâan 710127, China. Manuscript received September 09, 2023; revised August 26, 2015. and retrieving music files. In this context, automatic music genre classification has become increasingly important for organizing large audio collections. It involves analyzing dif- ferent musical characteristics, such as timbre, instrumentation, and lyrics [4]. Although classifying music genres is difficult because music is complex and varied, each genre usually has clear patterns. Good classification becomes possible when we properly model how the different features relate to each other. Several techniques for automatic music classification have been introduced [5], [6], but most were developed and eval- uated only on well-known Western music datasets. However, there is a great lack of research on the classification of Arabic music, especially when it comes to Yemeni music genres. Yemeni music ranks among the oldest and most culturally rich musical traditions in the Arab world. It comes from cen- turies of oral poetry, religious practices, and social traditions. This music shows Yemenâs rich and varied history, geography, and ethnic groups. Yemeni music can more accurately be seen as one of the main roots of Arab music [7]â[11]. A major challenge in classifying Arabic music, especially Yemeni music, is the serious shortage of available training data. In this paper, we have tackled this issue by generating the Yemeni Music Information Retrieval (YMIR) dataset, which encompasses data for the five primary genres: Sanaâani, Hadhrami, Tihami, Lahji, and Adeni. Additionally, we have proposed the Yemeni Music Classification Model (YMCM), a genre classification system built upon the widely recognized Convolutional Neural Network (CNN) architecture. We subse- quently carried out six main experimental groups, with each group including five sub-experiments. In the six main exper- iments, we used different time-frequency feature extraction methods: Mel-frequency Cepstral Coefficients (MFCCs) with 13, 20, and 40 coefficients (MFCC13, MFCC20, MFCC40), Mel spectrograms, Chroma features, and FilterBank features, all extracted from the YMIR dataset. For each feature type, we fed the extracted features separately into five different convolutional neural network models: our proposed YMCM, AlexNet, a standard CNN, VGG16, and MobileNet. This setup produced a total of 30 distinct experiments. The main contributions of this work are summarized as follows: âą We release YMIR, the first expert-annotated dataset for Yemeni music genre classification, comprising 1,475 audio clips across five traditional genres (Sanaâani, Hadhrami, Lahji, Tihami, and Adeni). The dataset was labeled by five Yemeni music experts using a struc- tured protocol, achieving strong inter-annotator agree- ment (Fleissâ Îș=0.85). arXiv:2604.05011v1 [cs.SD] 6 Apr 2026 JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20212 âą We propose the Yemeni Music Classification Model (YMCM), a CNN-based architecture with five convo- lutional layers designed for genre classification from timeâfrequency representations. âą We provide a consistent preprocessing and segmentation workflow and adopt stratified training/testing splits to support fair and repeatable evaluation. âą We systematically compare six of the feature extraction techniques: Mel-spectrograms, MFCCs with 13, 20, and 40 coefficients, Chroma, and FilterBank features, and quantify their effect on classification accuracy and sta- bility. The results indicate that Mel-spectrogram features realize the highest accuracy on the YMIR dataset when used with the proposed YMCM model, outperforming the other feature extraction techniques. âą We benchmark YMCM against established architectures (AlexNet, VGG16, MobileNet, and a baseline CNN) under identical settings, resulting in 30 experiments (six feature sets Ă five models). YMCM achieves the best overall consistent performance, reaching 98.83% accu- racy. I. PREVIOUS RESEARCH ON MUSIC GENRE CLASSIFICATION Music genre classification has been a well-explored area in Music Information Retrieval (MIR), with several techniques employing deep learning models for enhanced performance. In their seminal paper, Tzanetakis and Cook [5] define music genres as categorical labels that classify musical pieces based on elements such as instrumentation, rhythmic structure, and harmonic content. The authors identified three key features for analyzing musical content: timbral texture, rhythm, and pitch, particularly in the context of Western music styles like classical, jazz, pop, and rock. This foundational work has significantly influenced subsequent research in genre classi- fication. Both entire recordings and homogeneous segments within those recordings were utilized, achieving a classifica- tion accuracy of 61% across ten genres, using statistical pattern recognition classifiers. These results were closely aligned with those found in human genre classification studies. A notable contribution to this field is the work by Han Ding (2024), who proposed a novel hybrid model combining Resid- ual Networks (ResNet) with Bidirectional Gated Recurrent Units (Bi-GRU) for music genre classification. This approach leverages visual spectrograms as input, enabling the model to benefit from both the spatial feature extraction capabilities of ResNet and the temporal modeling capabilities of Bi- GRU. This method achieved promising results, showing the potential for deep learning models to significantly improve genre classification accuracy by capturing intricate patterns within the audio data [12]. In a similar vein, Oguike and Primus (2025) introduced a multimodal classification system for Sotho-Tswana mu- sical videos, incorporating audio, text (lyrics), and visual modalities. By using deep learning models for each modality and applying a decision-level fusion technique, their system demonstrated superior performance compared to unimodal models that rely solely on audio or lyrics [13]. Another innovative approach by Shen and Xiao (2024) applies Functional Data Analysis (FDA) to represent music signals as continuous functions, capturing both temporal and harmonic properties of music. This method, combined with Adaptive Fourier Decomposition (AFD), was tested on the GTZAN and FMA datasets, yielding significant improvements in classification accuracy over the traditional method [14]. Ahmed et al. (2024) explored the use of advanced deep learning models for music genre classification, comparing the effectiveness of various models such as CNNs, LSTMs, and SVMs. Their study highlights the superiority of CNNs in capturing complex spectrogram patterns, achieving high clas- sification accuracy on the GTZAN and ISMIR2004 datasets [15]. Beyond classical CNN baselines, recent studies have in- creasingly emphasized three practical levers for improving genre recognition: (1) augmentation and training-time regular- ization to mitigate limited dataset size, (2) systematic tuning of network hyperparameters, and (3) attention-based modeling for richer time-frequency context. For example, data aug- mentation coupled with deep architectures has been reported to substantially improve Mel-spectrogram-based classification on standard benchmarks [16]. Complementarily, automated configuration and hyperparameter optimization strategies have been explored to stabilize CNN performance across alternative spectral representations such as MFCC and STFT [17]. In a separate line of work, hybrid Transformer designs have been proposed to strengthen feature extraction from Mel- spectrograms by combining convolutional locality with global self-attention and channel-wise emphasis [18]. These trends collectively suggest that, for fair comparison on low-resource regional corpora, it is essential to control preprocessing, segmentation, and evaluation splits while benchmarking both feature representations and model capacity under identical experimental settings. This motivates the development of expert-labeled regional datasets and controlled baselines that make cross-model and cross-feature comparisons reproducible, which is the objective of the YMIR dataset and the proposed YMCM evaluation protocol. I. DESIGN OF YMIR This dataset is the first publicly accessible compilation of Arabic songs, with a particular focus on Yemeni music genres, offering a robust foundation for research into regional musical styles. The Yemeni Music Information Retrieval dataset encompasses data from the five main genres: Sanaâani, Hadhrami, Tihami, Lahji, and Adeni. Each musical scale also exhibits unique stylistic traits, making its identification closely related to genre classification in other types of music. Classification accuracy was assessed through inter-annotator agreement. Ultimately, the music recordings were labeled and organized to create the YMIR dataset. A. Data Collection The dataset comprises 1,475 audio clips across five tra- ditional Yemeni music genres: Sanaâani, Hadhrami, Lahji, Tihami, and Adeni. Each genre contains 295 audio files, with JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20213 Fig. 1. Waveform samples of the five YMIR genre classification dataset. all clips standardized to a duration of 30 seconds. This stan- dardization ensures a balanced representation across genres while also safeguarding the copyright of the original works. The YMIR dataset was compiled from online sources, including platforms such as YouTube. The audio files are provided in WAV format, with a sample rate of The YMIR dataset was compiled from online sources, including platforms such as YouTube. The audio files are provided in WAV format, with a sample rate of 48 kHz and a bit rate of 360 kbps. The dataset is organized into multiple folders, each corresponding to one of the five song genres. Accompanying metadata files are included, detailing each clipâs title, artist, genre, and file name. Fig. 1 shows the waveform representation of sample audio signals from the YMIR dataset. B. Data Labeling Experts in Yemeni music performed manual labeling to assign each audio track to its appropriate genre. Each track was labeled with a five-digit string, separated by underscores, representing its genre and position within the dataset. The string consists of the following elements: Song number, Sam- ple number, Song title, Artist name, and Genre (as illustrated in Fig. 2). This manual labeling process ensures the accuracy of classification tasks and facilitates the structured organization of the dataset. C. Judgments Five independent annotators, each possessing extensive ex- pertise in Yemeni music genres, contributed to the labeling process. The annotators were provided with comprehensive guidelines to ensure consistency and accuracy in their assess- ments. They were instructed to attentively listen to each audio clip and assign the most appropriate genre from the five core categories: Sanaâani, Hadhrami, Lahji, Tihami, and Adeni. Each annotator independently listened to all recordings and either assigned a genre or rejected the track if it did not clearly fit one of the five categories. In cases of disagreement, a consensus was reached through discussion, consultation, or majority voting. If three or more annotators agreed on a genre, the label was accepted for inclusion in the YMIR dataset; if not, it was rejected. For particularly ambiguous cases, annotators were advised to leave the track unannotated to avoid mislabeling. This process ensured consistency, accuracy, and reliability in the final labels. To evaluate the reliability of the dataset, we calculated Cohenâs Kappa score [19], a statistical measure that assesses inter-rater agreement for categorical classifications, given that five judges were involved in the labeling process. Îș = Ì P 0 â Ì P e 1â Ì P e (1) The factor 1â Ì P e how much the annotators agree with each other, beyond random guessing, while it Ì P 0 â Ì P e indicates how much real agreement there is, above and beyond random chance. A value of Îș = 1 signifies perfect agreement among all raters. Measuring agreement among the annotators for our dataset, as measured by Fleissâ kappa, yielded a score of 0.85, indicating a high level of agreement among the five raters. IV. METHODOLOGIES The methodology adopted in this study is described in this section, covering key components including data preprocess- ing steps, feature extraction techniques, and the foundational deep learning architectures serving as baseline models for the proposed music genre classification system. An overview of the studyâs complete workflow is illustrated in Fig. 3. Fig. 2. Data Labeling structure JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20214 Dataset Preprocessing Selected Models AlexNet CNN VGG 16 MobileNet Proposed YMCM Training Input Train Model Output Test the models Read 30 second WAV files Save the extracted data into a file Read the extracted from the file Feature extraction Extract FilterBanks Extract Mel-Spectrogram Extract Chroma Extract MFCC13 Extract MFCC20 Extract MFCC40 Segment Audio 5 segments Ă 6 seconds Evaluate the models Fig. 3. Workflow diagram of the proposed framework A. Data preprocessing In this study, all audio signals were processed using a unified preprocessing pipeline to ensure consistency across the dataset. Each audio file was resampled to 22,050 Hz and truncated to a fixed duration of 30 seconds. To han- dle the non-stationary characteristics of musical signals, a Short-Time Fourier Transform (STFT) is applied to obtain a timeâfrequency representation. The resulting power spectro- gram is computed once and subsequently reused to extract features, ensuring consistent spectral alignment across all representations while reducing redundant computations, as illustrated in Fig.4. To increase the number of training samples and capture temporal variations within each recording, every audio file was divided into five equal-length segments of 6 seconds each. After segmentation and preprocessing, the final dataset contained 7,258 samples, which were split into training and test sets using an 80:20 stratified split. This yielded 5,806 training and 1,452 testing samples, preserving the class distribution across both sets. B. Feature extraction The main purpose of feature extraction is to create a se- quence of feature vectors that offer a compact yet informative representation of the input audio signal. In music classification, feature extraction is a crucial step because the quality and relevance of the extracted features directly determine how accurate and effective the classification will be. Paying close attention to this step is essential, as it directly affects how well the following classification algorithms perform. As discussed earlier, music genre classification involves selecting suitable audio features and designing an effective Audio Signal STFT Output Fig. 4. WAV-format audio data Short-Time Fourier Transform Extraction. classification model. Earlier research on music genre clas- sification has mainly relied on four key feature types: Mel- spectrograms [20]â[23], FilterBanks [24], Chroma [25]â[27], and Mel Frequency Cepstral Coefficients (MFCC) [28]â[31]. In light of these approaches, we incorporated all four feature types in our experiments to evaluate their performance and determine which would yield the best results for the YMIR dataset. 1) Chroma: Chroma features are widely used in MIR [32], and are based on the twelve-tone equal temperament system. Because notes that are exactly one octave apart sound very similar to the human ear, the distribution of chroma features, which ignores the specific octave, still captures important musical information. This approach can highlight perceived similarities between musical elements that might not be obvious in the original frequency spectrum. Chroma features are usually appointed as a 12-dimensional vector v = [v(1),v(2),v(3),...,v(12)], where each com- ponent corresponds to one of the twelve pitch classes: JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20215 C,C#,D,D#,E,F,F #,G,G#,A,A#, and B. They show how the audio signalâs energy is distributed among the twelve different pitch classes (notes in the chromatic scale). 2) FilterBanks: A Mel filter bank consists of triangular filters specifically designed to approximate the way the human ear perceives differences in pitch and frequency. This design provides higher frequency resolution at lower frequencies and lower resolution at higher frequencies. Mel filter banks give more detail in low frequencies and less in high frequencies, which closely matches the way humans perceive sound [24]. 3) Mel-Spectrograms: The signal is divided into frames, and a Fast Fourier Transform (FFT) is computed for each frame. Subsequently, a Mel-scale is applied, dividing the entire frequency spectrum into uniformly spaced bands. A spectro- gram is then generated, where, for each frame, the signal magnitude is decomposed into its components corresponding to the frequencies in the Mel-scale. 4) Mel-Frequency Cepstral Coefficients (MFCCs): For fea- ture extraction in audio processing, MFCCs remain a standard choice across several sub-disciplines. Their utility is well- documented in the literature, ranging from the classification of music genres [28]â[31] and the detection of emotional cues in music [33] to the broader requirements of speech recognition [34]. By mimicking the non-linear frequency perception of the human ear, MFCCs provide a psychoacoustically motivated framework for feature extraction, which remains a standard approach in speech recognition research. Their alignment with human auditory perception further extends their utility to music analysis, where capturing psychoacoustic nuances is essential [35]. To capture the spectral characteristics of the signal, the MFCC pipeline first involves windowing the continuous waveform into discrete frames. A periodogram is then employed to estimate the power spectral density for each individual frame. C. Selected Classification Architecture Models Leveraging the robust feature-learning capabilities of CNN, prior studies have successfully deployed various architectures to categorize complex audio signals. In particular, AlexNet, VGG, and MobileNet have emerged as prominent choices for sound classification tasks. The following subsections offer a comparative overview of these architectures within the context of acoustic processing. âą CNN architectural foundations of modern deep learning were established with the introduction of LeNet-5. While originally developed for character recognition, this pio- neering framework introduced the essential concepts of local connectivity and shared weights through interleaved convolutional and subsampling layers. These principles proved vital for acoustic analysis, as they allow the network to achieve translation invariance, enabling the de- tection of specific sound patterns or pitch shifts regardless of their exact temporal position within a spectrogram. âą AlexNet [36] gained prominence following its decisive performance in the 2012 ImageNet Large Scale Vi- sual Recognition Challenge (ILSVRC). This milestone is widely regarded as an inflection point for the field, as it demonstrated the potential of deep convolutional ar- chitectures to outperform traditional hand-crafted feature methods in complex pattern recognition tasks. âą VGG networks, developed by the Visual Geometry Group at Oxford in 2014, VGG-style architectures shifted the paradigm toward the systematic use of small 3Ă 3 convolutional kernels. By stacking these filters in deep sequences, the design achieved a larger receptive field and increased model depth while simultaneously reducing the total number of parameters, a strategy that significantly enhanced its capacity for complex feature learning. âą MobileNet, Introduced by Howard et al. [37], MobileNet represents a specialized class of lightweight architectures engineered for deployment in resource-constrained envi- ronments. The core innovation lies in its use of depthwise separable convolutions, which decouple spatial filtering from feature generation. This architectural shift drasti- cally minimizes both parameter count and computational overhead, facilitating real-time sound classification on embedded hardware without significant loss in accuracy. D. Proposed YMCM Architecture Model The adoption of CNNs in signal processing and music information retrieval is largely driven by their capacity for hierarchical representation learning. Within this framework, initial layers are typically sensitive to fundamental acoustic attributes, such as localized spectral textures and transient energy fluctuations. As the architecture deepens, these primi- tive features are aggregated into higher-order abstractions that correspond to complex musical properties, including timbre, harmonic density, and cepstral envelopes. This inherent ability to learn multi-scale features provides a robust motivation for employing CNN-based models in musical genre classification. Drawing inspiration from the AlexNet framework, the pro- posed YMCM model utilizes a five-layer convolutional back- bone with a filter distribution of 64, 192, 384, 256, and 256, respectively. The initial feature extraction is performed by an 11Ă11 kernel with a stride of 4, transitioning to a 5Ă5 kernel in the second layer, and 3Ă3 kernels for all subsequent stages. To ensure stable convergence and introduce non-linearity, each convolution is coupled with batch normalization and ReLU activation. Spatial dimensionality reduction is achieved via max-pooling layers using a 3Ă 3 window. The architecture concludes with two dense layersâcomprising 1024 and 512 neuronsâleading to a final softmax-activated layer for genre categorization, as depicted in Fig. 5. V. EXPERIMENTS AND RESULTS This section details the datasets, training configurations, and performance metrics utilized to validate the proposed framework. The experimental design is structured to provide a comprehensive evaluation of the systemâs robustness and generalizability across diverse acoustic environments A. Training Protocol In this study, standard implementations of CNN, AlexNet, VGG16, and MobileNet were utilized with their original JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20216 Conv2D 64 filters, 11x11, strides 4 Conv2D 192 filters, 5x5, 1 Conv2D 384 filters, 3x3, 1 Conv2D 256 filters, 3x3, 1 Conv2D 256 filters, 3x3, 1 Normalized Flatten Dense 1024 unites ReLU Dense 512 unites ReLU Dropout(0.5) Dropout(0.5) softmax Output layer Fully connected layer Convolution layer Hidden layer Batch Normalization Max Pooling 2D Dropout(0.25) Fig. 5. Workflow diagram of the proposed framework network architectures and default parameter configurations. In contrast, the Yemeni Music Classification Model (YMCM) was implemented as a CNN, with architectural modifications. Audio feature extraction was performed using well- established Python-based libraries. Chroma, Mel Spectrogram, and MFCC features were extracted using the librosa library (version 0.11.0) [38], while FilterBank features were obtained using the Python Speech Features library (version 0.6). We derived Mel-spectrograms utilizing 128 discrete frequency bands, Chroma features with 12 pitch classes, and MFCC fea- tures with 13, 20, and 40 coefficients, following the standard configurations provided by the respective libraries. Each classification model was trained using a single fea- ture type per experimental configuration. For the experiments reported in this work, MFCC features with 13, 20, and 40 coefficients were adopted. The models were implemented using the Keras deep learning framework (version 2.9.0) with a TensorFlow 2.9.0 backend. All experiments were conducted on a workstation running Python 3.8 on Ubuntu 20.04, equipped with one virtual GPU (vGPU) with 32 GB of memory. GPU acceleration was enabled using the NVIDIA CUDA Toolkit version 11.2 to enhance computational efficiency during train- ing. The network was trained using the Adam optimizer with a fixed learning rate of 10 â4 . To minimize the discrepancy between predicted and ground-truth genre labels, we employed categorical cross-entropy as the objective function. The train- ing process was capped at 50 epochs with a mini-batch size of 16; however, to prevent overfitting and ensure optimal general- ization, an early stopping mechanism was implemented. This strategy monitored the validation loss and terminated training if no improvement was observed for 10 consecutive epochs. The specific hyperparameter configurations are consolidated in Table I. TABLE I TRAINING PARAMETERS USED IN THE EXPERIMENTS ParameterValue Number of Epochs50 Batch Size16 Learning Rate0.0001 OptimizerAdam Loss FunctionCategorical Crossentropy Early Stopping Patience10 epochs B. Evaluation Metrics The performance of the proposed Yemeni Music Classi- fication Model (YMCM) system is evaluated using several key metrics: weighted precision, sensitivity (Recall), F1-score, specificity, and balanced accuracy. These metrics, which assess various aspects of classification performance, were calculated as shown in the following formulas: Precision = TP TP + FP (1) Recall = TP TP + FN (2) F1-score = 2Ă PrecisionĂ Recall Precision + Recall (3) Specificity = TN FP + TN (4) JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20217 Features Extraction using 6 techniques A AlexNet 1-Chroma 2-FilterBanks 3-Mel-Spectrogram 4-MFCC13 5-MFCC20 6-MFCC40 B VGG 16 C Mobile Net D CNN E YMCM Test different models using different feature extraction techniques 1- Acc. 74.66% 2- Acc. 94.77% 3- Acc. 97.25% 4- Acc. 94.83% 5- Acc. 96.63% 6- Acc. 97.59% 1- Acc. 71.70% 2- Acc. 93.39% 3- Acc. 95.87% 4- Acc. 91.53% 5- Acc. 93.04% 6- Acc. 92.08% 1- Acc. 48.83% 2- Acc. 90.63% 3- Acc. 93.66% 4- Acc. 93.73% 5- Acc. 93.73% 6- Acc. 91.67% 1- Acc. 71.25% 2- Acc. 78.25% 3- Acc. 96.28% 4- Acc. 81.61% 5- Acc.94.41% 6- Acc. 92.84% 1- Acc. 82.85% 2- Acc. 96.28% 3- Acc. 98.83% 4- Acc. 97.59% 5- Acc. 97.18% 6- Acc. 97.52% Experiments Fig. 6. Summary of the experimental design and configurations conducted in this study. Accuracy = 1 n X TP FP + TN + TP + FN (5) The performance of the model is quantified using the fun- damental components of the confusion matrix: True Positives (TP), False Positives (FP), True Negatives (TN), and False Negatives (FN). In this context, TP and TN represent instances where the predicted labels align correctly with the ground truth for positive and negative classes, respectively. Conversely, FP and FN denote misclassifications, where the model erroneously assigns a positive label to a negative instance or fails to identify a true positive case. C. Experiment Steps As illustrated in Fig. 6, the experimental evaluation is organized into six main experimental groups, each comprising five sub-experiments. In the six main experiments, different timeâfrequency feature extraction techniques are employed, including MFCCs with 13, 20, and 40 coefficients (MFCC13, MFCC20, and MFCC40), Mel spectrograms, Chroma features, and FilterBank features, which are extracted from the YMIR dataset. For each feature representation, the extracted fea- tures are independently used as input to five CNN models, namely the proposed YMCM, AlexNet, CNN, VGG16, and MobileNet. This experimental configuration results in a total of 30 distinct experiments. D. Discussion and results An evaluation of the YMCM framework across various fea- ture extraction methodologies reveals a high degree of robust- ness, with the model maintaining consistently strong classifi- cation accuracy, as shown in Table I. When Chroma features are used, the model achieves an accuracy of 82.85%, which reflects the limited discriminative capacity of pitch-class in- formation that omits spectral energy distribution and temporal detail. The use of FilterBank features substantially improves performance to 96.28% accuracy by capturing perceptually motivated spectral energies; however, these representations do not explicitly preserve fine-grained timeâfrequency continuity. The highest performance is achieved with Mel-Spectrogram features, where YMCM attains up to 98.83% accuracy, pri- marily because Mel-Spectrograms preserve detailed spectral- temporal structures over time, enabling the model to jointly learn short-term dynamics, such as onsets and transients, as well as longer-term patterns related to rhythm and timbre. In contrast, MFCC-based representations achieve accuracies of 97.59%, 97.18%, and 97.52% for MFCC13, MFCC20, and MFCC40, respectively; however, their reliance on a discrete cosine transform compresses the spectral representation and partially removes frequency locality, which limits the ability of convolutional layers to exploit spatial correlations. Overall, these results demonstrate that while YMCM effectively lever- ages both compact cepstral and dense spectral representations, Mel-Spectrograms provide the most informative and discrimi- native features due to their preservation of spectral continuity and strong compatibility with convolutional feature learning. The validation loss and accuracy curves in Figs. 7 and 8 illustrate the training behavior of he YMCM model with different feature extraction techniques. Among all features, the Mel-Spectrogram consistently exhibits the most stable convergence, characterized by a rapid reduction in validation loss and a smooth increase in validation accuracy, ultimately achieving the highest accuracy of approximately 0.99. In contrast, FilterBanks and Chroma features show noticeable JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20218 01020304050 Epochs 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 Validation Accuracy 0.988 Validation Accuracy Comparison (Best Model Highlighted) Mel-Spectrogram (Best) Filterbanks Chroma MFCC-13 MFCC-20 MFCC-40 Fig. 7. Validation accuracy of the proposed YMCM framework using several feature extraction methods. 01020304050 Epochs 0 1 2 3 4 5 Validation Loss Validation Loss Comparison (Best Model Highlighted) Mel-Spectrogram (Best) Filterbanks Chroma MFCC-13 MFCC-20 MFCC-40 Fig. 8. Validation loss of the proposed YMCM framework using several feature extraction methods. fluctuations in both loss and accuracy, indicating less stable optimization and weaker generalization. MFCC-based features (MFCC13, MFCC20, and MFCC40) demonstrate improved convergence compared to Chroma and FilterBanks; however, they still exhibit slightly higher variance and slower stabiliza- tion than Mel-Spectrograms. The confusion matrices for each feature type, processed through the YMCM model, are presented in Fig. 9. The mel spectrogram features achieved the highest classification accu- racy, with strong diagonal dominance (290â291 correct pre- dictions per class) and minimal misclassification. Filterbanks and MFCC variants also performed competitively, though with slightly increased inter-class confusion. In contrast, chroma features exhibited the weakest performance, with significant off-diagonal errors particularly for Class 4 and Class 5, where only 212 and 226 samples were correctly classified, respec- tively. Based on the results reported in Table I, the optimal feature extraction strategy depends on the underlying model architecture. For AlexNet, MFCC40 achieves the highest accuracy 97.59%, indicating that higher-dimensional cep- stral representations effectively complement its large recep- tive fields. VGG16 and the baseline CNN attain their best performance using Mel-Spectrogram features, 95.87% and TABLE I PERFORMANCE RESULTS OF DIFFERENT MODELS USING VARIOUS FEATURE EXTRACTION TECHNIQUES. ModelFeatureAcc.Prec.Rec.F1. Extraction(%)(%)(%)(%) AlexNet Chroma74.6675.2074.6674.64 FilterBanks94.7794.7994.7794.76 Mel-Spectrogram97.2597.2797.2597.24 MFCC1394.8395.0394.8394.84 MFCC2096.6396.7696.6396.64 MFCC4097.5997.6197.5997.59 VGG16 Chroma71.0770.8771.0770.73 FilterBanks93.3993.4093.3993.39 Mel-Spectrogram95.8795.9095.8795.87 MFCC1391.5391.5791.5391.52 MFCC2093.0493.2093.0493.06 MFCC4092.0892.1292.0892.07 Chroma48.8347.6748.8347.49 FilterBanks90.6390.6990.6390.64 MobileMel-Spectrogram93.6693.7093.6693.65 NetMFCC1393.7393.7593.7393.73 MFCC2093.7393.7593.7393.72 MFCC4091.6791.7391.6791.67 CNN Chroma71.2562.0270.2565.02 FilterBanks78.2578.8968.2568.58 Mel-Spectrogram96.2896.5396.2896.32 MFCC1381.6185.6781.6181.89 MFCC2094.3594.4194.3594.35 MFCC4092.8493.4792.8492.90 Chroma82.8583.1582.8582.75 FilterBanks96.2896.3196.2896.27 ProposedMel-Spectrogram98.8398.8498.8398.83 YMCMMFCC1397.5997.6197.5997.59 MFCC2097.1897.2597.1897.19 MFCC4097.5297.5497.5297.52 96.28%, respectively, highlighting the importance of pre- serving detailed timeâfrequency structures for deeper con- volutional networks. For the lightweight MobileNet archi- tecture, MFCC13/MFCC20 provides the best trade-off be- tween accuracy 93.73% and computational efficiency. The proposed YMCM model achieves its highest accuracy with Mel-Spectrograms at 98.83%, confirming that dense spectral- temporal representations are particularly well suited to its design. Overall, these results indicate that Mel-Spectrograms are more effective for deeper and more expressive convo- lutional architectures, whereas MFCC-based features offer a more efficient alternative for lightweight models. Eventually, under identical experimental conditions, the YMCM model consistently outperforms all benchmark architectures listed. TABLE I OPTIMAL FEATURE EXTRACTION STRATEGY FOR EACH MODEL ModelOptimal FeatureAccuracy AlexNetMFCC4097.59% VGG16Mel-Spectrogram95.87% MobileNetMFCC13 / MFCC2093.73% CNN (Baseline)Mel-Spectrogram96.28% YMCM (Proposed)Mel-Spectrogram98.83% JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20219 (a) Filterbanks(b) Chroma(c) Mel Spectrogram (d) MFCC 13(e) MFCC 40(f) MFCC 20 Fig. 9. Confusion matrices for the proposed YMCM model with different audio feature extraction methods: (a) Filterbanks, (b) chroma features, (c) mel spectrogram, (d) 13 MFCCs, (e) 40 MFCCs, (f) 20 MFCCs. All matrices represent a 5-class classification task. These findings suggest that the YMCM model more effectively leverages complex audio representations, demonstrating a su- perior capacity for generalization. The consistency of these re- sults across varied acoustic conditions underscores the modelâs robustness for high-fidelity musical genre categorization. VI. CONCLUSION This paper introduced a comprehensive end-to-end frame- work specifically engineered for the classification of Yemeni musical genres, addressing the significant underrepresentation of culturally distinct Arabic traditions in the current Music Information Retrieval (MIR) landscape. Central to this con- tribution is the YMIR dataset, an expert-annotated corpus encompassing five foundational genres, validated through rig- orous inter-annotator agreement to ensure high label fidelity. By deploying the YMCM architecture, we demonstrate that a CNN-based approach tailored for time-frequency representa- tions can effectively capture the unique rhythmic and melodic nuances of Yemeni music within a standardized experimen- tal framework. Extensive experiments across multiple feature representations demonstrated that dense spectralâtemporal fea- tures are particularly effective for deep convolutional learning in this task. The proposed YMCM achieved the best overall performance, reaching 98.83% accuracy with Mel-spectrogram features and outperforming AlexNet, VGG16, MobileNet, and a baseline CNN under the same conditions. The results further indicate that feature selection is closely coupled with model capacity: Mel-spectrograms typically benefit deeper, more expressive CNNs, whereas MFCC variants remain competitive for lightweight architectures. Future work will extend YMIR in both scale and diversity (e.g., broader regional coverage and additional sub-genres) and investigate more advanced paradigms such as transformer- based audio encoders, self-supervised representation learning, and cross-dataset transfer learning to improve robustness in low-resource music classification settings. AI STATEMENT We used AI tools to improve grammar, readability, and wording. We did not use AI to generate data, analyses, figures, or conclusions. The authors reviewed and verified all content and take full responsibility for it. ACKNOWLEDGEMENT This work was supported by Yidan University Education Foundation under Grant JJA202507. REFERENCES [1] A. Elbir and N. Aydin, âMusic genre classification and music recom- mendation by using deep learning,â Electronics Letters, vol. 56, no. 12, p. 627â629, 2020. [2] J.-J. Aucouturier and F. Pachet, âRepresenting musical genre: A state of the art,â Journal of New Music Research, vol. 32, no. 1, p. 83â93, 2003. [3] A. Elbir, H. B. C ̧ am, M. E. Ì Iyican, B. Ì Ozt Ì urk, and N. Aydın, âMusic genre classification and recommendation by using machine learning techniques,â in 2018 Innovations in Intelligent Systems and Applications Conference (ASYU). IEEE, 2018, p. 1â5. JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 202110 [4] M. Serwach and B. Stasiak, âGa-based parameterization and feature selection for automatic music genre recognition,â in Proceedings of 2016 17th International Conference Computational Problems of Electrical Engineering (CPEE), 2016. [5] G. Tzanetakis and P. Cook, âMusical genre classification of audio signals,â IEEE Transactions on Speech and Audio Processing, vol. 10, no. 5, p. 293â302, 2002. [6] E. Wold, T. Blum, D. Keislar, and J. Wheaton, âContent-based classifi- cation, search, and retrieval of audio,â IEEE Multimedia, vol. 3, no. 3, p. 27â36, 1996. [7] J. Lambert, âA historical glimpse of music in yemen in the 1930s: The cylinders recorded by hans helfritz,â The World of Music (new series), vol. 9, no. 2, p. 49â66, 2020. [Online]. Available: https://journals.uni-goettingen.de/wom/article/view/1426 [8] G. Lavin, âMusic in colonial aden: Globalization, cultural politics and the record industry in an indian ocean port city, c. 1937â1960,â SOAS Musicology Series, 2020. [Online]. Available: https://w.academia. edu/73304097 [9] J. A. Ahmed, From the Yemeni Music Scene. Cultural Affairs Publishing House, 2007. [10] M. Altwaiji, âCultural heritage of yemeni folk song: The importance of singing,â International Journal of Sociology and Political Science, vol. 6, no. 1, 2024. [Online]. Available: https://w.sociologyjournal. in/assets/archives/2024/vol6issue1/6001.pdf [11] UNESCO,âAl-ghinaal-sanâani,traditional sanaa-stylesinging,âhttps://ich.unesco.org/en/RL/ al-ghina-al-sanani-traditional-sanaa-style-singing-00161, 2003. [12] H. Ding, L. Zhai, C. Zhao, F. Wang, G. Wang, W. Xi, Z. Wang, and J. Zhao, âGenre classification empowered by knowledge-embedded music representation,â IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, p. 2764â2776, 2024. [13] O. E. Oguike and M. Primus, âMultimodal music genre classification of sotho-tswana musical videos,â IEEE Access, 2025. [14] J. Shen and G. Xiao, âMusic genre classification based on functional data analysis,â IEEE Access, 2024. [15] M. Ahmed, U. Rozario, M. M. Kabir, Z. Aung, J. Shin, and M. Mridha, âMusical genre classification using advanced audio analysis and deep learning techniques,â IEEE Open Journal of the Computer Society, 2024. [16] T. C. Ba, T. D. T. Le, and L. T. Van, âMusic genre classification using deep neural networks and data augmentation,â Entertainment Computing, vol. 53, p. 100929, 2025. [17] T. Li, âOptimizing the configuration of deep learning models for music genre classification,â Heliyon, p. e24892, 2024. [18] P. Wu, W. Gao, Y. Chen, F. Xu, Y. Ji, J. Tu, and H. Lin, âAn improved vit model for music genre classification based on mel spectrogram,â PLOS ONE, vol. 20, no. 3, p. e0319027, 2025. [19] J. J. Randolph, âFree-marginal multirater kappa (multirater k [free]): An alternative to fleissâ fixed-marginal multirater kappa,â Online sub- mission, 2005. [20] S. Chillara, A. Kavitha, S. A. Neginhal, S. Haldia, and K. Vidyul- latha, âMusic genre classification using machine learning algorithms: a comparison,â International Research Journal of Engineering and Technology, vol. 6, no. 5, p. 851â858, 2019. [21] A. Dhall, Y. S. Murthy, and S. G. Koolagudi, âMusic genre classification with convolutional neural networks and comparison with f, q, and mel spectrogram-based images,â in Advances in Speech and Music Technology. Springer, 2021, p. 235â248. [22] H. Tang and N. Chen, âCombining cnn and broad learning for music classification,â IEICE Transactions on Information and Systems, vol. 103, no. 3, p. 695â701, 2020. [23] A. Ghildiyal, K. Singh, and S. Sharma, âMusic genre classification using machine learning,â in 2020 4th International Conference on Electronics, Communication and Aerospace Technology (ICECA). IEEE, 2020, p. 1368â1372. [24] H. Fayek, âSpeech processing for machine learning: Filter banks, mel-frequency cepstral coefficients (mfccs) and whatâs in-between,â URL:https://haythamfayek.com/2016/04/21/speech-processingfor- machine-learning. html, 2016. [25] X. Zhang, N. Li, and W. Li, âVerification for robustness of chroma feature,â Computer Science, p. S1, 2014. [26] M. Singh, S. K. Jha, B. Singh, and B. Rajput, âDeep learning neural networks for music information retrieval,â in 2021 International Confer- ence on Computational Intelligence and Knowledge Economy (ICCIKE). IEEE, 2021, p. 500â503. [27] S. Pulipati, C. P. Sai, K. S. Krishna, and C. Akhil, âMusic genre classification using convolutional neural networks,â Design Engineering, p. 2727â2732, 2021. [28] M. A. Kızrak, K. S. Bayram, and B. Bolat, âClassification of classic turkish music makams,â in 2014 IEEE International Symposium on In- novations in Intelligent Systems and Applications (INISTA) Proceedings. IEEE, 2014, p. 394â397. [29] R. Thiruvengatanadhan, âMusic genre classification using mfcc and aann,â International Research Journal of Engineering and Technology (IRJET), 2018. [30] P. Mandal, I. Nath, N. Gupta, K. Jha Madhav, G. Ganguly Dev, and S. Pal, âAutomatic music genre detection using artificial neural networks,â Intelligent Computing in Engineering, p. 17â24, 2020. [31] A. K. Sharma, G. Aggarwal, S. Bhardwaj, P. Chakrabarti, T. Chakrabarti, J. H. Abawajy et al., âClassification of indian classical music with time-series matching deep learning approach,â IEEE Access, vol. 9, p. 102 041â102 052, 2021. [32] M. A. Bartsch and G. H. Wakefield, âTo catch a chorus: Using chroma- based representations for audio thumbnailing,â in Proceedings of the 2001 IEEE Workshop on the Applications of Signal Processing to Audio and Acoustics (Cat. No. 01TH8575). IEEE, 2001, p. 15â18. [33] J. Dutta and D. Chanda, âMusic emotion recognition in assamese songs using mfcc features and mlp classifier,â in 2021 International Conference on Intelligent Technologies (CONIT). IEEE, 2021, p. 1â5. [34] A. X. Glittas, L. Gopalakrishnan et al., âA low latency modular- level deeply integrated mfcc feature extraction architecture for speech recognition,â Integration, vol. 76, p. 69â75, 2021. [35] T. L. Li and A. B. Chan, âGenre classification and the invariance of mfcc features to key and tempo,â in International Conference on MultiMedia Modeling. Springer, 2011, p. 317â327. [36] A. Krizhevsky, I. Sutskever, and G. E. Hinton, âImagenet classification with deep convolutional neural networks,â Advances in neural informa- tion processing systems, vol. 25, 2012. [37] A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, âMobilenets: Efficient convo- lutional neural networks for mobile vision applications,â arXiv preprint arXiv:1704.04861, 2017. [38] B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto, âlibrosa: Audio and music signal analysis in python.â SciPy, vol. 2015, p. 18â24, 2015.