Paper deep dive
TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios
Hong Lyu, Mingru Yang, Qianhua He, Yanxiong Li, Jinxin Huang, Zhengyu Pei
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/8/2026, 5:11:36 AM
Summary
The paper introduces the TriA Pipeline, a large-scale automatic audio annotation system designed to convert raw audio from diverse streaming platforms into high-quality training data with event annotations. The pipeline comprises four stages: standardization, audio activity detection, audio event detection, and filtering. Using this pipeline, the authors construct the TriA dataset containing over 2130 hours of audio across 431 classes. A prior-knowledge-guided subset (TriA_GK) is derived and evaluated on three domestic audio classification tasks (DESED, Kitchen20, Nonspeech7k). Experimental results demonstrate that combining TriA_GK with manually annotated data yields average relative improvements of 3.97% in accuracy and 3.35% in Macro-F1, validating the pipeline's effectiveness in addressing data scarcity in specific acoustic scenarios.
Entities (12)
Relation Signals (11)
TriA Pipeline → develops → TriA Dataset
confidence 97% · A TriA dataset was constructed with the TriA Pipeline, over 2130 hours of audio covering 431 audio classes.
Audio Event Detection → uses → BEATs
confidence 96% · the BEATs model is employed to annotate audio events for AAD segments.
TriA Pipeline → consistsof → Audio Activity Detection
confidence 95% · TriA Pipeline consists of four main stages: Standardization, Audio Activity Detection (AAD), Audio Event Detection (AED), and Filtering.
TriA Pipeline → consistsof → Filtering
confidence 95% · TriA Pipeline consists of four main stages: Standardization, Audio Activity Detection (AAD), Audio Event Detection (AED), and Filtering.
TriA Pipeline → consistsof → Audio Event Detection
confidence 95% · TriA Pipeline consists of four main stages: Standardization, Audio Activity Detection (AAD), Audio Event Detection (AED), and Filtering.
TriA Pipeline → consistsof → Standardization
confidence 95% · TriA Pipeline consists of four main stages: Standardization, Audio Activity Detection (AAD), Audio Event Detection (AED), and Filtering.
TriA Dataset → contains → TriA_GK
confidence 95% · we partitioned a prior-knowledge-guided subset (TriA_GK) from TriA
Filtering → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:There are some datasets of varying scales for audio classification (AC) applied to different tasks. However, annotated data is limited for most scenarios, such as domestic environments. To address this challenge, we propose an $\textbf{A}$utomatic $\textbf{A}$udio $\textbf{A}$nnotation Pipeline--TriA Pipeline, which can efficiently convert audio from various scenarios into high-quality training data with audio event annotations. A TriA dataset was constructed with the TriA Pipeline, over 2130 hours of audio covering 431 audio classes. Furthermore, we partitioned a prior-knowledge-guided subset (TriA$_{\mathrm{GK}}$) from TriA and conduct comparative experiments on three domestic AC tasks. Comparing the result on manually annotated data only and that on manually annotated data combines TriA$_{\mathrm{GK}}$, TriA$_{\mathrm{GK}}$ could achieve average relative gains of 3.97% in accuracy and 3.35% in Macro-F1, validating the effectiveness of TriA$_{\mathrm{GK}}$ and the TriA Pipeline.
Tags
Links
- Source: https://arxiv.org/abs/2607.06179v1
- Canonical: https://arxiv.org/abs/2607.06179v1
Trouble viewing inline? Open PDF directly →
Full Text
30,428 characters extracted from source content.
Expand or collapse full text
TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios Hong Lyu ID , Mingru Yang ID , Qianhua He ID ∗ , Yanxiong Li ID ∗ , Jinxin Huang, Zhengyu Pei School of Electronic and Information Engineering, South China University of Technology, Guangzhou, China eehlyu@mail.scut.edu.cn, eeqhhe@scut.edu.cn, eeyxli@scut.edu.cn Abstract There are some datasets of varying scales for audio classifica- tion (AC) applied to different tasks. However, annotated data is limited for most scenarios, such as domestic environments. To address this challenge, we propose an Automatic Audio Annotation Pipeline–TriA Pipeline, which can efficiently con- vert audio from various scenarios into high-quality training data with audio event annotations. A TriA dataset was constructed with the TriA Pipeline, over 2130 hours of audio covering 431 audio classes. Furthermore, we partitioned a prior-knowledge- guided subset (TriA GK ) from TriA and conduct comparative ex- periments on three domestic AC tasks. Comparing the result on manually annotated data only and that on manually annotated data combines TriA GK , TriA GK could achieve average relative gains of 3.97% in accuracy and 3.35% in Macro-F1, validating the effectiveness of TriA GK and the TriA Pipeline. Index Terms: Automatic audio annotation pipeline, Audio dataset, Audio classification 1. Introduction Audio Classification (AC) enables the recognition of various en- vironmental sound events, serving as a core component in appli- cations ranging from multimedia content analysis [1] and audio captioning [2] to bio-acoustic monitoring [3, 4]. Ultimately, the efficacy of these AC systems heavily depends on the availability of large-scale audio datasets that cover diverse acoustic scenes. Existing AC datasets can be broadly categorized into general-purpose (e.g., AudioSet [1], FSD50K [5], ESC-50 [6], FSC-89 [7]) and specialized ones (e.g., Kitchen20 [8], CHiMe- Home [9], NSynth-100 [10]). The former encompass a wide variety of acoustic scenes, while the latter focus on specific do- mains, such as domestic environments or human non-speech sounds. Among general-purpose AC datasets, AudioSet is the most extensive, consisting of a class-balanced subset (AS-20K), a class-unbalanced subset (AS-2M), and an evaluation set, with AS-2M being widely utilized for pre-training and fine-tuning audio models [11, 12, 13, 14]. FSD50K is another large-scale audio dataset, but remains class-unbalanced, leaving certain sound classes underrepresented, such as wails, moans, wheezes, squeals, and specific domestic sounds. ESC-50, by contrast, is a widely used class-balanced dataset of 50 classes across 5 ma- jor categories, though its scale is notably limited. In summary, existing general-purpose AC datasets suffer from two primary limitations: (i) insufficient data for specific acoustic scenes, and (i) limited dataset scale. Among specialized AC datasets, DESED [15] comprises real recordings (DESEDreal) and synthetic data, covering 10 ** indicates the corresponding author. classes of domestic audio events, and is widely used for AC and sound event detection (SED) in domestic scenes [16, 17, 18]. However, DESED real contains only 5955 annotated clips. Kitchen20 is designed for kitchen AC tasks, while both HTAD [19] and CHiMe-Home target domestic activity recog- nition. CIRDO [20] and BiMP [21] are simulated datasets for safety monitoring of elderly individuals living alone, and Non- speech7k [22], originally developed for paralinguistic classifi- cation, can similarly be applied to domestic safety monitoring. Nevertheless, all these specialized datasets suffer from limited scale. In general, annotated audio data remains scarce across many scenarios, with the problem becoming more pronounced in highly specific domains. To address the scarcity of annotated audio data in spe- cific scenarios, we propose a large-scale Automatic Audio Annotation Pipeline–TriA Pipeline. TriA Pipeline consists of four stages: Standardization, Audio Activity Detection (AAD), Audio Event Detection (AED), and Filtering, which collec- tively convert raw audio from diverse streaming platforms and scenarios into high-quality training data with audio event an- notations.The most closely related works are Emilia-Pipe [23], NVSpeech-Pipe [24], and NonVerbalSpeech-Pipe [25]. However, Emilia-Pipe relies on Automatic Speech Recognition (ASR) and supports only speech data annotation, while both NVSpeech-Pipe and NonVerbalSpeech-Pipe focus exclusively on paralinguistic speech annotation. In contrast, by integrating an AED module, TriA Pipeline supports audio event annota- tion across a broad range of scenarios, resolving the annotation scarcity problem in the specific domains described above. Both objective metrics and subjective listening evaluations confirm that TriA Pipeline produces high-quality and diverse audio data with high annotation reliability. Based on TriA Pipeline, the TriA dataset is further con- structed, containing over 2130 hours of high-quality audio data covering 431 audio classes. To evaluate the effectiveness of TriA, we partition a prior-knowledge-guided subset, TriA GK , from TriA and setup three specific classification tasks: DESED for Audio Classification (DESED AC ), Kitchen20, and Non- speech7k. They represent three domestic AC tasks: general do- mestic AC, kitchen AC, and domestic safety monitoring. We conduct comparative experiments using TriA GK and manually annotated datasets. Experimental results show that fine-tuning models using only TriA GK achieves performance comparable to models fine-tuned using manually labeled data. Further- more, fine-tuning the model by combining TriA GK with manu- ally annotated data can achieve average relative improvements of 3.97% in accuracy and 3.35% in Macro-F1. These results indicate that TriA GK can help the model achieve better perfor- mance in the specified classification task, and further validate the effectiveness of the proposed TriA Pipeline. arXiv:2607.06179v1 [eess.AS] 7 Jul 2026 StandardizationStandardization Audio Activity Detection Audio Activity Detection Audio Event Detection Audio Event Detection FilteringFiltering AADsegmentsAED segments [Speech, Music][Speech] mp3mp3 mp3mp3 Overview of the TriAPipeline Figure 1: Overview of the TriA Pipeline. AAD and AED denote Audio Activity Detection and Audio Event Detection. 2. TriA Pipeline This section details the TriA Pipeline and its evaluation on mini- batch data. As illustrated in Figure 1, the TriA Pipeline consists of four main stages: Standardization, Audio Activity Detection (AAD), Audio Event Detection (AED), and Filtering. 2.1. Standardization This step follows the same procedure as Emilia-Pipe and aims to standardize audio with heterogeneous formats for subsequent processing. Specifically, the original audio recordings are con- verted into mono-channel WAV format with a sample rate of 24 kHz and a 16-bit sample width. The target loudness is then nor- malized to -20 dBFS, while the signal amplitude is constrained within the range of -3 dB to 3 dB to prevent distortion. Finally, each waveform is normalized by dividing all sample points by the maximum amplitude. 2.2. Audio Activity Detection The purpose of this step is to split long audio recordings into clips of appropriate length while removing redundant segments. The auditok tool 1 is used to split each sample into AAD seg- ments based on a preset energy threshold, with the maximum segment duration constrained to 30 s. To determine the appropriate minimum duration for sound activity segments and maximum duration for silent segments, subjective listening tests can be conducted on datasets related to specific scenarios. For each audio class in each dataset, 10 instances are randomly selected for listening. During the tests, the minimum duration required to identify the event in each in- stance and the temporal interval between two adjacent events of the same class are recorded. Based on the test results, the average minimum time required to identify events of each class is calculated and referred to as the Event Critical Time (ECT). The average interval between adjacent events of each class is also calculated and referred to as the Silent Critical Time (SCT). During the AAD processing, retained activity segments are re- quired to exceed the minimum ECT, as shorter segments con- tain insufficient information to identify. The retained silent seg- ments are restricted to be shorter than the maximum SCT, since excessively long silent segments introduce redundant informa- tion and reduce efficiency. In our experiments, the target sce- nario is the domestic scene, where the minimum ECT and the maximum SCT are 1.2 s and 2.0 s, respectively. 2.3. Audio Event Detection To enable the TriA dataset to be directly applied to AC tasks, the BEATs model 2 is employed to annotate audio events for AAD segments. The BEATs model achieves state-of-the-art (SOTA) performance on the AS-2M dataset under the single au- 1 https://github.com/amsehili/auditok 2 https://github.com/microsoft/unilm/tree/ master/beats dio modality and can detect 527 audio classes [12]. Specifically, for each AAD segment, the AS-2M fine-tuned BEATs iter3+ model is used for event detection, dividing it into multiple segments with event annotations. To detect short events in the segment, local detection is conducted by scanning each AAD segment with a fixed detection window length and window shift, producing preliminary detection segments. Ad- jacent detection segments within the same AAD segment are then concatenated if their annotated Top-1 events are identi- cal. To further detect long audio in the segment and continu- ous event across adjacent detection windows, global detection is performed on each concatenated detection segment. In global detection, the window length is equal to the length of the seg- ment to be detected. If the globally detected class differs from the original class, the newly class is appended to the segment annotation. The resulting segments are called AED segments. The window length for AED local detection is required to exceed the maximum ECT (3 s). The purpose is to enable the model to process one or more target events as completely as possible in a single detection, providing sufficient information to the model and improving the reliability of the detection re- sults. However, if the window length is too long, the audio seg- ment detected by the model at one time may include too many events, which will bring difficulties to the model detection. The higher the event confidence threshold for AED detection, the higher the reliability of the detection results. In our experiment, the event confidence threshold is set to 0.6. The window length and window shift for local detection are set to 5 s and 3 s, re- spectively. 2.4. Filtering The original audio may exhibit varying quality and the BEATs model may produce detection errors. To ensure the data quality and improve the matching degree between data and event anno- tations, the audiobox-aesthetics 3 and the CLAP model 4 are used to filter AED segments. Specifically, Production Complexity (PC) and Production Quality (PQ) in aesthetics are used as fil- tering indicators. PC and PQ are relatively objective indicators, focusing on the complexity of the audio scene and the technical quality respectively [26]. For each AED segment, the PC and PQ, as well as the CLAP similarity between the segment and event annotation [27], are calculated. Segments with PC, PQ, or CLAP similarity below predefined thresholds are filtered. The remaining high-quality segments are stored in MP3 format, and a JSONL (JSON Lines) file containing the associated metadata is generated for efficient indexing and retrieval. 2.5. Evaluation on minibatch data To validate the effectiveness of each individual module and the overall TriA Pipeline, 284.7 hours of original audio are ran- 3 https://github.com/facebookresearch/ audiobox-aesthetics 4 https://github.com/microsoft/CLAP Table 1: Statistical results of 285 hours of original data processed by the TriA Pipeline. Filtering 1 uses PC, PQ, and CLAP similarity thresholds of 1.8, 5.5, and 5, respectively, while Filtering 2 uses thresholds of 2.24, 5.85, and 7.43, respectively. DataDuration (s)Aesthetics PC| PQCLAP SimilarityTotal Duration (hours) Total Classes minmaxavg± stdminavg± stdminavg± std Original Audio2.0417605.37233.05± 841.701.39| 3.125.17± 1.66| 7.20± 0.89—284.70 (100.00%)— Annotated w/o Filtering1.2030.0010.87± 9.221.33| 2.854.60± 1.71| 6.52± 1.07-4.197.71± 2.52242.62 (85.22%)— Annotated w Filtering 1 1.2030.0011.57± 9.381.80| 5.504.90± 1.62| 6.92± 0.772.008.04± 2.24182.37 (64.06%)325 Annotated w Filtering 2 1.2030.0011.70± 9.832.24| 5.855.18± 1.45| 7.10± 0.717.439.54± 1.6580.08 (28.13%)258 domly collected from streaming platforms and processed using the pipeline. The experiment is conducted on an NVIDIA RTX 3090 GPU, with a total processing duration of about 10 hours and the RTF of 0.03. As shown in Table 1, statistics are per- formed across multiple objective evaluation perspectives. The original audio exhibits a broad duration distribution, with a rel- atively long average duration and a large variance. In contrast, the filtered data shows a more standardized and appropriate duration. Compared with the unfiltered data, the filtered data achieves much higher average PC, PQ, and CLAP Similarity, indicating that the data quality and annotation reliability have been significantly improved. However, the quality of the unfil- tered data is lower than that of the original data. The reason is that the original data has longer average duration and higher average sampling rate, which favor PC and PQ measures. Af- ter increasing the filtering threshold, the quality and annotation reliability are further improved. However, it reduces the class diversity and the total duration of the data. In addition, subjec- tive listening tests are also conducted, randomly selecting 100 samples to assess annotation accuracy, and an average accuracy of 93.67% is obtained. Overall, the results demonstrate that the TriA Pipeline is feasible and effective. It can increase the amount of data in each class and continuously expand the scale of training data through repeated large-scale audio data collec- tion and pipeline processing. 3. TriA Dataset 3.1. Statistics and Analysis Over 8706 hours of original audio data were collected from streaming platforms including Bilibili and Douyin, covering di- verse topics such as daily life, entertainment media, and tech- nology. After processing with the TriA Pipeline, the TriA dataset was constructed. The TriA dataset contains over 2130 hours of audio data, covering 431 audio classes. As illustrated in Figure 2, the class distribution of TriA is presented. Among them, the class with the largest number of samples is Music, which is attributed to the widespread presence of background music in streaming videos. Although the class distribution of the data is imbalanced, task-specific balanced subsets can be constructed for downstream applications. To analyze the data quality and annotation reliability of the TriA dataset, Table 2 summarizes statistics on PC, PQ, and CLAP similarity of different datasets. TriA achieves the best PC and PQ, indicating that the data quality of TriA is high, which benefits from the filtering stage in the TriA Pipeline. The CLAP similarity of TriA ranks third among all datasets, showing that its annotation reliability is not as good as that of Kitchen20 and Nonspeech7k, but better than that of DESED real . This differ- ence can be attributed to the annotation method of the datasets. Kitchen20 and Nonspeech7k are manually annotated. TriA is automatically annotated using the pipeline, whereas DESED real is annotated through crowdsourcing. TRIAL MODE − Click here for more information Figure 2: Class distribution of TriA. Table 2: PC, PQ, and CLAP similarity of different datasets. The metrics for TriA are calculated on a randomly sampled 300- hour subset. DataAesthetics PC| PQCLAP Similarity DESED real 3.39± 0.86| 5.85± 0.857.43± 4.89 Kitchen202.49± 0.37| 6.38± 0.7412.87± 3.41 Nonspeech7k2.24± 0.67| 6.34± 0.8011.71± 3.14 TriA5.13± 1.48| 7.08± 0.679.62± 1.54 3.2. Prior-knowledge-guided Subset The prior-knowledge refers to the threshold used for specific scenario, which was determined through subjective listening tests and statistics related to the specific scenario. The sam- ples related to these prior-guided thresholds are collected to a subset, TriA GK . For convenience of subsequent experiment, TriA GK is fur- ther partitioned into three subsets: TriA DESED , TriA Kitchen20 , and TriA Nonspeech7k . The audio classes they cover are the same as the classes of DESED real , Kitchen20, and Nonspeech7k respectively.The class nomenclature of the TriA follows the AudioSet ontology [1].For example, the class Elec- tric shavertoothbrush in DESED real corresponds to Electric shaver, electric razor and Electric toothbrush in TriA. Table 3 reports the number of clips and total duration for each dataset. The scale of each subset of TriA GK is larger than that of the corresponding dataset. 4. Experiments This section evaluates the effectiveness of TriA GK for three do- mestic AC tasks: DESED AC , Kitchen20, and Nonspeech7k. For each task, three experiments are conducted. The baseline model is trained with different datasets, the first on the manu- ally annotated data, the second on the TriA GK , and the third on both, the TriA GK first and then the manually annotated data (re- Table 3: Number of clips and total duration of different datasets. Unlabeled data are excluded from the statistics. DatasetClipsTotal Durations (hours) Manual annotate DESED real 595516.50 Kitchen2010701.47 Nonspeech7k70146.75 TriA GK TriA DESED 2351742.60 TriA Kitchen20 16882.45 TriA Nonspeech7k 961811.29 fer to Table 4), to analyze whether the pipeline-processed data can help the model learn the specified AC task. 4.1. Implementation Details 4.1.1. Baseline The BEATs model [12] is adopted. The backbone consists of a 12-layer Transformer encoder with 90M parameters, initialized with the pre-trained weights BEATs iter3+ . A linear classifica- tion head, comprising two fully connected layers, is appended to the backbone. Input audio is resampled to 16 kHz. We extract 128-dimensional Mel-filter bank features using a Povey window of 25 ms and a hop size of 10 ms. All experiments are conducted on an RTX 3090 GPU. The cross entropy loss and the AdamW optimizer are used. The maximum training epoch is set to 50, with an early stopping patience of 15. For the DESED AC and Nonspeech7k tasks, the entire BEATs model is fully fine-tuned. For the Kitchen20 task, only the classification head is fine-tuned. The learning rates for full fine-tuning and adapter fine-tuning are set to 5e-5 and 6e-3, respectively, following a cosine annealing learning rate schedule. Accuracy and Macro-F1 are used as evaluation metrics. Accuracy measures the overall classification correctness, while Macro-F1 reflects the consistency of model performance across different classes. The validation metric is accuracy + Macro-F1. 4.1.2. Task Setup The DESED real training set contains 4429 clips, consisting of the real strongly labeled set and the real weakly labeled set. The validation set and test set contain 373 and 1153 clips re- spectively, which are derived from the real strongly labeled validation set and the real strongly labeled test set. To adapt DESED real to the AC setting, the strongly labeled sets are con- verted into weak labels by retaining only the audio paths and event annotations. TriA DESED is split into training and valida- tion sets with a 9:1 ratio. The test set for the DESED AC task is the DESED real test set. Previous research evaluates Kitchen20 using 5-fold cross- validation. To unify the test set for the Kitchen20 task, the fifth fold of Kitchen20 dataset is used as the test set, which contains 160 clips. The training and validation sets of Kitchen20 dataset contain 480 and 160 clips, corresponding to the first three folds and the fourth fold, respectively. TriA Kitchen20 is divided into training and validation sets with a 4:1 ratio. The Nonspeech7k dataset provides 6289 training and 725 test clips. We further split the training clips into training and validation sets with a 9:1 ratio. TriA Nonspeech7k is also split into training and validation sets at a 9:1 ratio. The test set for the Nonspeech7k task is the Nonspeech7k test set. Table 4: Experimental results on three AC tasks. A + B denotes sequential fine-tuning, where the model is first fine-tuned on A and then further fine-tuned on B. TaskDatasetAccF1 DESED real 0.78370.7943 DESED AC TriA DESED 0.82550.7810 TriA DESED + DESED real 0.82580.8256 Kitchen200.92500.9272 Kitchen20TriA Kitchen20 0.93750.9355 TriA Kitchen20 + Kitchen200.98130.9812 Nonspeech7k0.94480.9437 Nonspeech7kTriA Nonspeech7k 0.89380.8734 TriA Nonspeech7k + Nonspeech7k0.94900.9464 4.2. Results Table 4 reports the experimental results on three AC tasks. The model fine-tuned on TriA DESED achieves higher accuracy than that fine-tuned on DESED real , but obtaining a lower Macro- F1. It suggests that TriA DESED provides better overall data quality but exhibits slightly lower consistency across different classes. When applying sequential fine-tuning (TriA DESED → DESED real ), the model significantly outperforms the one fine- tuned solely on DESED real , achieving relative gains of 5.37% in accuracy and 3.94% in F1. It indicates that TriA DESED helps the pre-trained model acquire task-relevant knowledge, and then DESED real further improves the performance of the model. TriA Kitchen20 outperforms Kitchen20 in both overall data quality and consistency across different audio classes. Sequen- tial fine-tuning (TriA Kitchen20 → Kitchen20) achieves relative improvements of 6.09% in accuracy and 5.82% in Macro-F1 compared to fine-tuning on Kitchen20 alone, demonstrating that TriA Kitchen20 can help the model to learn the Kitchen20 task. TriA Nonspeech7k performs worse than Nonspeech7k in both accuracy and Macro-F1.However, sequential fine-tuning (TriA Nonspeech7k → Nonspeech7k) still leads to performance gains over fine-tuning only on Nonspeech7k, with relative gains of 0.44% in accuracy and 0.29% in F1. It further demonstrates that TriA GK can help the model learn the specific AC tasks. Overall, fine-tuning on pipeline-processed data achieves performance comparable to, and in some cases better than, that obtained using manually annotated data.Compared to fine-tuning on manually annotated data alone, combining TriA GK subsets with manual data through sequential fine- tuning achieves average relative improvements of 3.97% in ac- curacy and 3.35% in Macro-F1. It validates the effectiveness of TriA GK for domestic AC tasks and further confirms the feasi- bility and practical value of the TriA Pipeline. 5. Conclusion In this paper, we propose the TriA Pipeline, a large-scale automatic audio annotation pipeline.It efficiently converts audio collected from various streaming platforms into high- quality training data with event annotations. Based on the TriA Pipeline, the TriA dataset is constructed, which contains over 2130 hours of audio data covering 431 audio classes. Through comparative experiments, we verify the effectiveness of the prior-knowledge-guided subset, TriA GK , on domestic AC tasks, and further confirm that the TriA Pipeline is effective. The TriA Pipeline code and dataset is now released 5 . 5 https://github.com/huanxian/TriA 6. Acknowledgments This work was partly supported by the national natural science foundation of China (62371195, 62111530145), and the ex- change project of the 10th Meeting of China-Croatia Science and Technology Cooperation Committee (10-34). 7. Generative AI Use Disclosure We used GPT-5.2 to assist in polishing the manuscript. 8. References [1] J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE inter- national conference on acoustics, speech and signal processing (ICASSP), 2017, p. 776–780. [2] K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio cap- tioning dataset,” in ICASSP 2020-2020 IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP), 2020, p. 736–740. [3] R. R. Kvsn, J. Montgomery, S. Garg, and M. Charleston, “Bioa- coustics data analysis–a taxonomy, survey and open challenges,” IEEE Access, vol. 8, p. 57 684–57 708, 2020. [4] A. Terenzi, N. Ortolani, I. Nolasco, E. Benetos, and S. Cecchi, “Comparison of feature extraction methods for sound-based clas- sification of honey bee activity,” IEEE/ACM transactions on au- dio, speech, and language processing, vol. 30, p. 112–122, 2021. [5] E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra, “Fsd50k: an open dataset of human-labeled sound events,” IEEE/ACM Trans- actions on Audio, Speech, and Language Processing, vol. 30, p. 829–852, 2021. [6] K. J. Piczak, “ESC: Dataset for Environmental Sound Classi- fication,” in Proceedings of the 23rd Annual ACM Conference on Multimedia.ACM Press, 2015, p. 1015–1018. [Online]. Available: http://dl.acm.org/citation.cfm?doid=2733373.2806390 [7] Y. Li, W. Cao, J. Tan, Q. Li, and G. Chen, “Few-shot class-incremental audio classification using pseudo-incrementally trained embedding learner and continually updated stochastic classifier,” IEEE Transactions on Audio, Speech and Language Processing, vol. 33, p. 3880–3895, 2025. [8] M. Moreaux, M. G. Ortiz, I. Ferran ́ e, and F. Lerasle, “Benchmark for kitchen20, a daily life dataset for audio-based human action recognition,” in 2019 International Conference on Content-Based Multimedia Indexing (CBMI), 2019, p. 1–6. [9] P. Foster, S. Sigtia, S. Krstulovic, J. Barker, and M. D. Plumbley, “Chime-home: A dataset for sound source recognition in a do- mestic environment,” in 2015 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2015, p. 1–5. [10] Y. Li, J. Tan, Q. Li, G. Chen, S. Huang, and T. Virtanen, “Few-shot open-set audio classification using attention information-fused prototypes,” IEEE Transactions on Audio, Speech and Language Processing, vol. 34, p. 1929–1943, 2026. [11] Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumb- ley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, p. 2880–2894, 2020. [12] S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “Beats: audio pre-training with acoustic tok- enizers,” in Proceedings of the 40th International Conference on Machine Learning, 2023, p. 5178–5193. [13] X. LI and X. Li, “Atst: Audio representation learning with teacher-student transformer,” in Proc. Interspeech 2022, 2022, p. 4172–4176. [14] H. Dinkel, Z. Yan, Y. Wang, J. Zhang, Y. Wang, and B. Wang, “Streaming audio transformers for online audio tagging,” in Proc. Interspeech 2024, 2024, p. 1145–1149. [15] N. Turpault, R. Serizel, A. Parag Shah, and J. Salamon, “Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,” in Workshop on Detection and Classification of Acoustic Scenes and Events, 2019. [16] Z. Lin, Y. Li, Z. Huang, W. Zhang, Y. Tan, Y. Chen, and Q. He, “Domestic activities clustering from audio recordings using con- volutional capsule autoencoder network,” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, p. 835–839. [17] T. Khandelwal and R. K. Das, “A Multi-Task Learning Frame- work for Sound Event Detection using High-level Acoustic Char- acteristics of Sounds,” in Interspeech 2023, 2023, p. 1214–1218. [18] T. Khandelwal, R. K. Das, A. Koh, and E. S. Chng, “Leverag- ing audio-tagging assisted sound event detection using weakified strong labels and frequency dynamic convolutions,” in 2023 IEEE Statistical Signal Processing Workshop (SSP), 2023, p. 329–333. [19] E. Garcia-Ceja, V. Thambawita, S. A. Hicks, D. Jha, P. Jakob- sen, H. L. Hammer, P. Halvorsen, and M. A. Riegler, “Htad: A home-tasks activities dataset with wrist-accelerometer and audio features,” in International Conference on Multimedia Modeling. Springer, 2021, p. 196–205. [20] M. Vacher, S. Bouakaz, M.-E. Bobillier-Chaumon, F. Aman, R. A. Khan, S. Bekkadja, F. Portet, E. Guillou, S. Rossato, and B. Lecouteux, “The cirdo corpus: comprehensive audio/video database of domestic falls of elderly people,” in 10th Interna- tional Conference on Language Resources and Evaluation, 2016, p. 1389–1396. [21] J. Dibble and M. C. Bazzocchi, “Bi-modal multiperspective per- cussive (bimp) dataset for visual and audio human fall detection,” IEEE Access, 2025. [22] M. M. Rashid, G. Li, and C. Du, “Nonspeech7k dataset: Clas- sification and analysis of human non-speech sound,” IET Signal Processing, vol. 17, no. 6, p. e12233, 2023. [23] H. He, Z. Shang, C. Wang, X. Li, Y. Gu, H. Hua, L. Liu, C. Yang, J. Li, P. Shi et al., “Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,” in IEEE Spoken Language Technology Workshop (SLT), 2024, p. 885–890. [24] H. Liao, Q. Ni, Y. Wang, Y. Lu, H. Zhan, P. Xie, Q. Zhang, and Z. Wu, “Nvspeech: An integrated and scalable pipeline for human-like speech modeling with paralinguistic vocalizations,” arXiv preprint arXiv:2508.04195, 2025. [25] R. Ye, Y. Zhou, R. Yu, Z. Lin, K. Li, X. Li, X. Liu, G. Zeng, and Z. Wu, “A scalable pipeline for enabling non-verbal speech generation and understanding,” arXiv preprint arXiv:2508.05385, 2025. [26] A. Tjandra, Y.-C. Wu, B. Guo, J. Hoffman, B. Ellis, A. Vyas, B. Shi, S. Chen, M. Le, N. Zacharov, C. Wood, A. Lee, and W.-N. Hsu, “Meta audiobox aesthetics: Unified automatic qual- ity assessment for speech, music, and sound,” arXiv preprint arXiv:2502.05139, 2025. [27] B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang, “Clap learning audio concepts from natural language supervision,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, p. 1–5.