Paper deep dive
RealDESED: A Real-World Domestic Sound Event Detection Benchmark
Florian Schmid, Paul Primus, Alexander Fichtinger, Tara Jadidi, Tobias Morocutti, Gerhard Widmer
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/21/2026, 3:58:20 AM
Summary
The paper introduces RealDESED, a real-world domestic sound event detection (SED) benchmark consisting of 5,710 audio recordings collected from 652 participants in their homes. Unlike existing benchmarks that rely on simulated soundscapes or web-crawled data, RealDESED captures natural domestic environments with realistic variability in devices, placement, and acoustic conditions. The dataset features 15 domestic sound classes, multi-annotator temporal labels, and rich metadata. The authors establish a transformer-based baseline (ATST-F) achieving a macro-averaged PSDS1 score of 0.731, and investigate annotation aggregation, post-processing, and long-form inference strategies.
Entities (12)
Relation Signals (10)
RealDESED â contains â 5,710 audio recordings
confidence 99% · RealDESED, a real-world domestic sound event detection (SED) benchmark comprising 5,710 audio recordings collected by 652 participants in their homes.
RealDESED â achievesscore â 0.731
confidence 95% · Our baseline achieves a macro-averaged PSDS1 score of 0.731 on the test set.
Gerhard Widmer â affiliatedwith â LIT Artificial Intelligence Lab
confidence 95% · Gerhard Widmer 1,2 ... 2 LIT Artificial Intelligence Lab, Linz, Austria
Florian Schmid â affiliatedwith â Johannes Kepler University Linz
confidence 95% · Florian Schmid 1â ... 1 Institute of Computational Perception, Johannes Kepler University Linz, Austria
RealDESED â collectedby â 652 participants
confidence 95% · comprising 5,710 audio recordings collected by 652 participants in their homes.
RealDESED â hassoundclasses â 15 common domestic sound classes
confidence 95% · contains temporally precise annotations for 15 common domestic sound classes.
ATST-F â isbaselinefor â RealDESED
confidence 90% · We establish a strong transformer-based baseline using ATST-F [20]... fine-tuned on RealDESED.
ATST-F â pretrainedon â AudioSet-Strong
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This paper presents RealDESED, a real-world domestic sound event detection (SED) benchmark comprising 5,710 audio recordings collected by 652 participants in their homes. Each recording is between 15 and 35 seconds long and contains temporally precise annotations for 15 common domestic sound classes. In contrast to existing SED datasets, which typically rely on simulated soundscapes or broad web-crawled audio, RealDESED consists exclusively of recordings captured in natural domestic environments, reflecting realistic variability in recording devices, device placement, acoustic conditions, background sounds, and naturally occurring event co-occurrences. A distinguishing characteristic of the dataset is its multi-annotator labeling scheme, where each recording is independently annotated by multiple annotators, while the validation and test sets undergo an additional review process to ensure high annotation quality and reliable benchmarking. Furthermore, the dataset provides rich metadata, including recording device, device placement, environment labels, and textual scene descriptions. We establish a strong transformer-based baseline and investigate annotation aggregation strategies, post-processing methods, long-form inference, and the impact of recording metadata on model performance. Our baseline achieves a macro-averaged PSDS1 score of 0.731 on the test set. We believe RealDESED provides a valuable benchmark for developing and evaluating robust SED systems under realistic domestic conditions, helping to bridge the gap between current research benchmarks and real-world deployment.
Tags
Links
- Source: https://arxiv.org/abs/2607.16736v1
- Canonical: https://arxiv.org/abs/2607.16736v1
Trouble viewing inline? Open PDF directly â
Full Text
32,238 characters extracted from source content.
Expand or collapse full text
Detection and Classification of Acoustic Scenes and Events 202628â29 October 2026, Boston, MA, USA RealDESED: A Real-World Domestic Sound Event Detection Benchmark Florian Schmid 1â , Paul Primus 1â , Alexander Fichtinger 1 , Tara Jadidi 1 , Tobias Morocutti 1 , Gerhard Widmer 1,2 1 Institute of Computational Perception, Johannes Kepler University Linz, Austria 2 LIT Artificial Intelligence Lab, Linz, Austria first.last@jku.at AbstractâThis paper presents RealDESED, a real-world domestic sound event detection (SED) benchmark comprising 5,710 audio recordings collected by 652 participants in their homes. Each recording is between 15 and 35 seconds long and contains temporally precise annotations for 15 common domestic sound classes. In contrast to existing SED datasets, which typically rely on simulated soundscapes or broad web-crawled audio, RealDESED consists exclusively of recordings captured in natural domestic environments, reflecting realistic variability in recording devices, device placement, acoustic conditions, background sounds, and naturally occurring event co-occurrences. A distinguishing characteristic of the dataset is its multi-annotator labeling scheme, where each recording is independently annotated by multiple annotators, while the validation and test sets undergo an additional review process to ensure high annotation quality and reliable benchmarking. Furthermore, the dataset provides rich metadata, including recording device, device placement, environment labels, and textual scene descriptions. We establish a strong transformer- based baseline and investigate annotation aggregation strategies, post- processing methods, long-form inference, and the impact of recording metadata on model performance. Our baseline achieves a macro-averaged PSDS1 score of 0.731 on the test set. We believe RealDESED provides a valuable benchmark for developing and evaluating robust SED systems under realistic domestic conditions, helping to bridge the gap between current research benchmarks and real-world deployment. Index Termsâsound event detection, domestic environment recordings, real-world dataset, multi-annotator labeling, temporal event localization 1. INTRODUCTION The goal of Sound Event Detection (SED) is to identify acoustic events and their temporal boundaries within audio recordings. Reliable SED systems enable applications such as security and surveillance [1], smart homes and assistive technologies [2], [3], and health moni- toring [4]. Advances in deep learning have substantially improved SED performance over the past decade, making SED an important component of intelligent systems that interpret acoustic environments. Existing SED Benchmarks: Progress in SED has been driven by publicly available datasets and benchmarks, enabling reproducible development and evaluation. The largest strongly annotated benchmark is AudioSet Strong [5], [6], comprising about 100k web-crawled record- ings spanning 447 sound classes. Other SED datasets are typically an order of magnitude smaller and focus on specific domains. Among these, DESED [7], [8] is the dominant benchmark for domestic SED, combining synthetic training soundscapes with strongly annotated web- crawled recordings for evaluation. Similarly, UrbanSED [9] consists of synthetic urban soundscapes with automatically derived strong labels, while MAESTRO Real [10] provides long-form recordings with strong labels inferred from human-annotated weak labels, limiting temporal annotation granularity. Current Limitations: Existing SED datasets have several limi- tations. First, obtaining temporally precise annotations is expensive, resulting in relatively small strongly annotated datasets. For example, â These authors contributed equally to this work. The LIT AI Lab is supported by the Federal State of Upper Austria. Gerhard Widmerâs work is supported by the European Research Council under the European Unionâs Horizon 2020 research and innovation programme (Grant No. 101019375, Whither Music?). while weakly labeled AudioSet contains about 2M recordings, its strongly annotated subset is around 20 times smaller. Second, temporal annotations are inherently subjective, particularly regarding event continuity and onset and offset boundaries, motivating multiple independent annotations for robust label aggregation. A common alternative is synthetic soundscape generation with automatically derived strong labels, as used in DESED and UrbanSED. However, this introduces a domain gap between synthetic and real recordings, often reducing performance on real-world audio. Finally, many datasets rely on web-crawled recordings and fixed 10-second clips, limiting realism and temporal context. Research Directions in SED: These limitations have strongly influenced SED research. To reduce the reliance on expensive strong annotations, researchers have explored learning from weak labels [7], [11], [12], semi-supervised learning [7], heterogeneous supervision [13]â[15], and varying pre-training paradigms [16]â [19]. Together, these approaches improve data efficiency but remain constrained by the limited availability of realistic strongly annotated benchmark datasets. RealDESED: In this paper, we present RealDESED, a new dataset for domestic sound event detection that addresses the limitations of existing benchmarks. Compared with existing datasets, RealDESED offers the following key characteristics: âą Real-world domestic recordings collected by 652 participants in their homes, capturing realistic variability in recording devices, environments, and recording conditions, instead of relying on web-crawled recordings or simulated audio mixtures. âąLonger recordings of 15â35 seconds, providing longer acoustic context and compound action sequences than the fixed 10-second clips commonly used in existing SED benchmarks. âąA benchmark of 5,710 real recordings with predefined training, validation, and test splits, enabling system development and evaluation on real in-distribution data. âąMulti-annotator temporal annotations, with reviewed valida- tion and test labels. âąRich recording metadata, including recording device, device placement, environment labels, and scene descriptions. We establish a strong transformer-based baseline using ATST-F [20], pre-trained on AudioSet Strong [16] and fine-tuned on RealDESED. We further investigate annotation aggregation, long-form inference on recordings exceeding the modelâs input length, and the impact of recording metadata on system performance, establishing a com- prehensive benchmark for future research. To facilitate reproducible research, we publicly release the dataset, baseline implementation, and data collection and annotation guidelines. 1 2. DATASET CONSTRUCTION RealDESED was collected in the context of a university machine learning course. The dataset comprises realistic domestic sound 1 https://github.com/fschmid56/RealDESED https://zenodo.org/records/20056072 arXiv:2607.16736v1 [eess.AS] 18 Jul 2026 Detection and Classification of Acoustic Scenes and Events 202628â29 October 2026, Boston, MA, USA scenes with temporally precise annotations for 15 common domestic sound classes: bell ringing, coffee machine, cutlery and dishes, door opening or closing, footsteps, keyboard typing, keychain, light switch, microwave, phone ringing or notification, running water, toilet flushing, vacuum cleaner, wardrobe or drawer opening or closing, and window opening or closing. Dataset creation consisted of three stages: data collection, annotation, and review, followed by a benchmark split into training, validation, and test sets. 2.1. Data Collection Each participant was asked to collect eight to ten realistic domestic sound scenes between 15 and 35 seconds long. Recordings were made on consumer devices (smartphones, tablets, portable microphones, etc.) to reflect realistic home conditions. Submitted files were converted to a unified waveform format during preprocessing. The collection protocol encouraged balanced target-class coverage, multi-event and polyphonic scenes, and realistic acoustic conditions while avoiding unrealistically complex recordings. Natural background and non-target sounds, such as traffic noise, appliance hum, and chair movement were allowed and encouraged, provided the target events remained clearly audible. To increase acoustic diversity, collectors recorded scenes using static and mobile device placement, with the device placed on a surface or carried by the participant, respectively. Recordings containing speech were not permitted. Each recording was accompanied by structured metadata, including clip-level labels for audible target and non-target sounds, recording device, device placement, recording environment, and a short free- text scene description summarizing the sequence of actions and sound events. Non-target sounds, recording devices, and recording environments were not restricted to fixed vocabularies, whereas device placement was specified as either static or mobile. The metadata was later used for analysis and to support the annotation process. Collectors selected either C0 or C-BY as the license for their recordings. 2.2. Data Annotation The recordings were annotated by 645 annotators using Label Studio Enterprise [21]. Each annotator was assigned at least 20 clips from a pool of 5,710 audio files. The assignment included the annotatorâs own recordings when available, supplemented with randomly selected clips from other collectors. Annotators used an interface displaying the audio waveform, time- aligned spectrogram, and user-provided metadata. The metadata, including scene descriptions and target-class tags, was provided only as contextual guidance because it could be incomplete or inaccurate. For each recording, annotators identified all audible events belong- ing to the 15 target classes and assigned each event a class label and temporal region. Overlapping events were allowed, while repeated events of the same class were annotated as separate regions if separated by a perceptually distinct pause. Non-target events were ignored. In total, the annotation procedure produced 64,430 submitted annotations, with multiple annotations available for most recordings. 2.3. Annotation Review The review process served as a qualitative quality-control step for the annotations. Each reviewer was assigned five to ten annotated recordings, and each reviewed recording had at least two independent annotations. This allowed reviewers to compare annotator decisions, identify inconsistencies, and correct the annotations where necessary. For each recording, reviewers inspected the audio, metadata, waveform, spectrogram, and submitted annotations to verify class labels, temporal boundaries, and annotation completeness. When necessary, they corrected class labels and temporal regions. Exact boundary agreement across all annotations was not required, as small onset and offset differences may reflect perceptual ambiguity. In a subsequent quality-control step, approximately 300 low-quality reviews were identified based on remaining disagreement between annotations and reassigned for review, further improving consistency. 2.4. Data Split The collected dataset contains 5,710 audio files, partitioned into training, validation, and test sets by collector ID, ensuring all recordings from the same domestic environment were assigned to the same split. The training, validation, and test sets contain 3,704, 999, and 1,007 recordings, respectively. All validation and test recordings were reviewed according to the procedure outlined in Section 2.3, while approximately one third of the training set contains reviewed annotations. 3. DATASET CHARACTERISTICS The following section summarizes RealDESED, including the char- acteristics of its recordings, metadata, and temporal annotations, as well as the effect of the annotation review. 3.1. Collected Audio Recordings and Metadata The 5,710 recordings total 37.85 hours. Scene durations range from 12.03 to 40.64 seconds, averaging 23.86 seconds. Despite a few deviations from the 15â35-second target, most recordings follow the intended short domestic-scene format. 3.1.1. Target sounds: The creator-reported clip-level labels indicate that recordings contain between 1 and 5 target sounds, with an average of 2.80. The most frequent labels are footsteps (2,686 files) and door open close (1,791). The least frequent labels are bell ringing (430) and coffee machine (312). Common metadata-level co-occurrences include door open close with footsteps (1,278 files) and cutlery dishes with running water (629). Less plausible combinations, such as coffee machine with toilet flushing (1), are rare. 3.1.2. Non-target sounds and scene descriptions: The metadata shows substantial non-target sound diversity: 2,849 recordings (49.89%) contain at least one non-target label, spanning 1,734 normalized labels. Files contain 0.72 non-target labels on average, ranging from 0 to 8. The most frequent normalized non-target sounds are clothing rustle (215 occurrences) and chair movement (142). Free-text scene descriptions have an average length of 24.56 words, suggesting that most collectors provided informative descriptions. 3.1.3. Recording environments: We normalize 161 recording- environment strings into 65 categories. Kitchens are most common (1,686 recordings), followed by bedrooms (1,091) and hallways (1,071), while balconies, garages, and staircases appear only once. 3.1.4. Recording devices: The 523 raw recording-device strings were mapped to 342 device tags of varying specificity (âmobile phoneâ to âFairphone 5â). Most recordings were made with mobile phones, accounting for 5,242 files (91.80%). Tablets, laptops, and other devices were less common, accounting for 173 (3.03%), 172 (3.01%), and 123 recordings (2.15%), respectively. The metadata further indicates 3,086 recordings with static and 2,624 recordings with mobile device placement. 3.2. Annotations Before review, the collected dataset contains 64,430 non-aggregated events across all 5,710 files with a total length of 77.80 hours. On average, each file has 2.45 annotators who marked an average of 4.60 event regions per file. When taking the union of annotated regions within each file, 77.97% (29.51 hours) of the total readable audio duration are covered by at least one event annotation. Detection and Classification of Acoustic Scenes and Events 202628â29 October 2026, Boston, MA, USA 3.2.1. Class distribution: The distribution of the annotated events is imbalanced. The most frequent class is footsteps, with 11,597 raw annotation events, followed by door open close (7,978). The least frequent classes are coffee machine (1,140) and vacuum cleaner (1,278). This pattern is consistent with the metadata target class tags. 3.2.2. Agreement between target class tags and strong annotations: Metadata target classes are well covered by temporal annotations: 15,726 of 16,008 fileâtarget class pairs appear in at least one raw annotation for the same file (98.24%). Annotators also marked classes that were not listed in the metadata target tags. 2,641 of 64,430 annotated events (4.99%) are outside the metadata target tags, affecting 1,134 files (19.86%). The most common added classes are footsteps (533) and door open close (282). 3.2.3. Annotator agreement: For files with multiple annotators, class-set agreement is relatively high but not perfect. When ignoring temporal annotations and comparing only the set of classes marked as present in each file (similar to clip-level tags) the mean pairwise Jaccard similarity is 0.875. For cases where multiple annotators marked the same class in the same file, pairwise temporal agreement is also reasonably high. The mean intersection-over-union over full recordings is 0.694. The mean overlap divided by the shorter total annotated duration is 0.929, showing that annotators often mark overlapping regions even when their exact onset and offset boundaries differ. 3.2.4. Overlap analysis: At least one temporal overlap of different target classes occurs in 7,155 of 14,009 fileâannotator pairs (51.07%). Across individual annotators, 19.25% of annotated time is covered by two or more simultaneous events, with up to 4 simultaneous events. The most common overlapping class pairs reflect plausible domestic situations such as cutlery dishes with running water (974 fileâannotator pairs) and keyboard typing with phone ringing (866). 3.3. Review The validation and test splits were completely reviewed, covering all 999 validation files and all 1,007 test files. In the training split, 1,108 of 3,704 files were reviewed, corresponding to 29.91% of the training data. The review process substantially improves agreement between annotations: For reviewed annotations, the mean class-set Jaccard similarity increases from 0.875 to 0.994, meaning that after review most annotations agree exactly on the set of classes present in a file. The pairwise mean temporal intersection-over-union increases from 0.694 for raw unreviewed same-file same-class annotations to 0.862 for accepted reviewed annotations. Similarly, the mean overlap over the shorter annotation increases from 0.929 to 0.975. 4. BENCHMARK EXPERIMENTS We establish the first benchmark results on RealDESED using ATST- F [20], a frame-level Audio Spectrogram Transformer pre-trained on AudioSet Strong [16] and fine-tuned on RealDESED. Beyond establishing a strong baseline, we investigate three aspects specific to RealDESED: annotation aggregation for multi-annotator labels, long-form inference on recordings exceeding the modelâs input length, and the influence of recording metadata on SED performance. 4.1. Experimental Setup Training is performed on randomly sampled 10 s crops from the training recordings, with frame-level labels generated at 25 Hz. Since recordings exceed the model input length, different crops are sampled across epochs. Input preprocessing follows the original ATST-F setup [20]. The model is optimized using AdamW [22] with a peak learning rate of4Ă10 â5 , cosine learning-rate scheduling, Mixup [23], frequency warping [20], and binary cross-entropy loss. Table 1: Label aggregation strategies. Mean± std over three runs. # MethodPSDS1-MâPSDS1âPSDS2-Mâ 1 Random (Fixed).675± .003 .503± .010 .916± .002 2 Random (Epoch).692± .004 .521± .002 .933± .001 3 Majority.688± .004 .518± .004 .924± .002 4 Intersection.681± .007 .511± .012 .904± .001 5 Union.673± .004 .497± .002 .925± .002 6 Collector.682± .002 .507± .002 .920± .003 7 Uniform Soft.690± .004 .519± .005 .933± .002 8 Weighted Soft (α=16) .696± .007 .532± .007 .930± .003 9 Majority + Reviewed .696± .006 .533± .009 .927± .001 10 Weighted + Reviewed .697± .005 .533± .004 .929± .004 For long-form inference, recordings are processed using overlapping 10 s windows with a default hop size of 5 s. Unless stated otherwise, predictions are merged using triangular weighting (floor 0.3) and post-processed using a median filter with a window size of 360 ms. These settings are varied in the corresponding experiments below. Performance is evaluated using PSDS1 and PSDS2 scores [24]. We report standard PSDS1 score (α CT = 0,α ST = 1) and macro- averaged PSDS1 and PSDS2 scores. 4.2. Annotation Aggregation RealDESED provides multiple annotations per recording, while approximately one third of the training set additionally contains reviewed annotations. We compare different annotation aggregation strategies during training (Table 1). Random (Fixed) (Row #1) establishes the single-annotator baseline by selecting one annotation per recording before training. In contrast, Random (Epoch) (#2) samples one annotation per recording in every epoch, substantially improving performance across all metrics. Simple aggregation methods, including majority vote (#3), intersection (#4), and union (#5), consistently underperform this strategy. Row #6 uses only annotations from the recording collector. Although this outperforms the fixed single-annotator baseline (#1), it remains inferior to the best multi-annotator strategies. We next replace hard label aggregation with soft frame-wise targets (#7 and #8). Lety i â0, 1 CĂT denote the binary frame-wise labels of annotatori,Cis the number of sound classes,Tthe number of frames, andNthe number of annotators. For clarity, we omit the recording index throughout the following equations. Uniform soft labels are obtained by averaging the frame-wise annotations over the Nannotators, preserving annotator disagreement instead of collapsing it into hard labels. As shown in Row #7, this approach outperforms hard aggregation strategies and achieves the highest PSDS2 Macro score across all strategies. A natural extension is to weight annotators according to their annotation quality. We compute the frame-wise macro Dice agreement between pairs of annotators (i and j), D(y i ,y j ) = 1 |C âČ | X câC âČ 2âšy (c) i ,y (c) j â© â„y (c) i â„ 1 +â„y (c) j â„ 1 ,(1) whereC âČ denotes the set of active classes,âšÂ·,·â©the inner product, andâ„·℠1 theâ 1 -norm. For each annotator, their resulting Dice scores are averaged over all other annotators and recordings to obtain an annotator quality score s i . Final soft labels are then computed as w i = s α i P N j=1 s α j , Ìy = N X i=1 w i y i ,(2) whereα = 16is selected on the validation set. As shown in Row #8, quality-weighted soft labels consistently outperform uniform Detection and Classification of Acoustic Scenes and Events 202628â29 October 2026, Boston, MA, USA Table 2: Comparison of post-processing methods. Results are averaged over three independent runs. OverallClass-wise PSDS1 MethodPSDS1-M PSDS1 PSDS2-MBellCoffee Cutlery Door Steps Typing Keychain Light Micro Phone Water Toilet Vacuum Ward. Window Raw0.6620.4870.8750.749 0.8090.5440.406 0.461 0.8720.6270.398 0.716 0.721 0.875 0.7810.8970.4550.623 Median Filter0.6930.5280.9030.748 0.8270.5940.465 0.565 0.8660.6590.404 0.787 0.751 0.903 0.7730.9490.4710.631 cSEBBs0.7310.5730.9250.789 0.8610.6370.512 0.584 0.9020.7060.442 0.784 0.768 0.925 0.8600.9680.5300.692 Table 3: Influence of long-form inference parameters. CategorySettingPSDS1-MâPSDS1âPSDS2-Mâ Aggregation Average.731.573.925 Maximum.721.565.910 Filter floor 0.0.703.551.881 0.3.731.573.925 0.6.730.573.917 1.0.726.569.915 Hop size (s) 2.5.736.579.919 5.0.731.573.925 7.5.718.561.910 10.0.711.553.904 averaging and achieve the best PSDS1 scores among methods without incorporating manually reviewed annotations. Finally, Rows #9 and #10 evaluate the benefit of incorporating the reviewed training annotations available for one third of the training split. The reviewed annotations replace the corresponding raw annotations, while the remaining recordings continue to use majority voting or weighted soft labels (based on #3 and #8), respectively. Reviewed annotations yield a larger PSDS1 gain for majority voting than for weighted soft labels. The comparatively small improvement for the latter suggests that annotator-aware weighting already produces higher-quality targets, making this automatic aggregation strategy competitive with human review. Overall, effectively leveraging multi- annotator annotations improves PSDS1 Macro from .675 to .697. 4.3. Post-Processing Frame-wise SED predictions require post-processing to obtain tem- porally consistent event predictions. Building on the best label aggregation strategy (Row #10 in Table 1) in terms of PSDS1 score, which uses a 360 ms median filter, we optimize the post-processing strategy. We compare raw predictions, per-class median filtering, and cSEBBs [25], with hyperparameters optimized on the validation set. Table 2 summarizes the results. Both post-processing methods sub- stantially outperform the raw predictions, highlighting the importance of temporal modeling. Median filtering improves the PSDS1 Macro score from .662 to .693, but class-wise optimization provides no additional benefit and tends to overfit the validation set. cSEBBs achieves the best overall performance with .731 and is therefore used in all subsequent experiments. Class-wise performance varies considerably across sound events. Sustained events such as vacuum cleaner or running water are detected reliably, whereas short-duration events including light switch or door open close remain the most challenging. cSEBBs improves performance for nearly all classes, demonstrating the benefit of strong temporal modeling. 4.4. Long-Form Modeling Since RealDESED recordings exceed the 10 s input length of the baseline model, inference is performed using overlapping sliding windows. Table 3 compares different window aggregation methods, triangular weighting functions, and hop sizes. Average aggregation consistently outperforms maximum aggrega- tion. Likewise, triangular weighting improves over uniform averaging Fig. 1: Macro-averaged PSDS1 over three runs grouped by recording device, device placement, and recording environment. (floor = 1.0), with a floor of 0.3 achieving the best overall performance. Finally, smaller hop sizes improve detection performance by increasing overlap between neighboring windows, although the gain from reducing the hop size from 5 s to 2.5 s is modest relative to the additional computational cost. We therefore use average aggregation, a triangular filter floor of 0.3, and a hop size of 5 s. 4.5. Metadata Analysis A key advantage of RealDESED is the availability of rich recording metadata, enabling analyses of model performance across different recording conditions. To illustrate the potential of the provided metadata, Figure 1 reports macro-averaged PSDS1 grouped by recording device, device placement, and recording environment. For the device analysis, we compare Android smartphones against Apple devices (iPhones and iPads), excluding heterogeneous device categories such as laptops and other recording hardware. The baseline exhibits meaningful performance differences across all three metadata dimensions. iOS devices outperform Android devices, mobile recordings outperform static recordings, and performance also varies considerably across recording environments. These results demonstrate that recording conditions can have a substantial impact on SED performance, motivating more detailed investigations into factors such as device characteristics, recording setup, and environment- specific event distributions. Such analyses highlight the potential of RealDESED as a benchmark for metadata-aware learning, domain adaptation, and device-robust SED. 5. CONCLUSION We presented RealDESED, a new benchmark for domestic sound event detection comprising 5,710 real-world recordings collected by 652 participants in their homes. Compared to existing SED datasets, RealDESED provides recordings captured directly in target envi- ronments, temporally precise multi-annotator annotations, reviewed evaluation labels, and rich recording metadata. We established a strong transformer-based baseline and demonstrated the importance of annotation quality, temporal post-processing, and long-form inference, while highlighting the influence of recording conditions on SED perfor- mance. By publicly releasing the dataset, baseline implementation, and complete data collection and annotation protocol, we hope to provide a valuable resource for the community and a realistic benchmark for developing robust sound event detection systems that bridge the gap between current SED benchmarks and real-world deployment. Detection and Classification of Acoustic Scenes and Events 202628â29 October 2026, Boston, MA, USA REFERENCES [1] R. Radhakrishnan, A. Divakaran, and A. Smaragdis, âAudio analysis for surveillance applications,â in Proc. WASPAA. IEEE, 2005, p. 158â161. [2]C. Debes, A. Merentitis, S. Sukhanov, M. E. Niessen, N. Frangiadakis, and A. Bauer, âMonitoring activities of daily living in smart homes: Understanding human behavior,â IEEE Signal Process. Mag., vol. 33, no. 2, p. 81â94, 2016. [3] R. M. Alsina-Pag ` es, J. Navarro, F. Al Ì Ä±as, and M. Herv Ì as, âhomesound: Real-time audio event detection based on high performance computing for behaviour and surveillance remote monitoring,â Sensors, vol. 17, no. 4, p. 854, 2017. [4] Y. Zigel, D. Litvak, and I. Gannot, âA method for automatic fall detection of elderly people using floor vibrations and sound - proof of concept on human mimicking doll falls,â IEEE Trans. Biomed. Eng., vol. 56, no. 12, p. 2858â2867, 2009. [5] J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, âAudio set: An ontology and human-labeled dataset for audio events,â in Proc. ICASSP, 2017, p. 776â780. [6] S. Hershey, D. P. W. Ellis, E. Fonseca, A. Jansen, C. Liu, R. C. Moore, and M. Plakal, âThe benefit of temporally-strong labels in audio event classification,â in Proc. ICASSP, 2021, p. 366â370. [7]N. Turpault, R. Serizel, J. Salamon, and A. P. Shah, âSound event detection in domestic environments with weakly labeled data and soundscape synthesis,â in Proc. DCASE, 2019, p. 253â257. [8]R. Serizel, N. Turpault, A. P. Shah, and J. Salamon, âSound event detection in synthetic domestic environments,â in Proc. ICASSP, 2020, p. 86â90. [9]J. Salamon, D. MacConnell, M. Cartwright, P. Li, and J. P. Bello, âScaper: A library for soundscape synthesis and augmentation,â in Proc. WASPAA, 2017, p. 344â348. [10]I. Mart Ì Ä±n-Morat Ì o and A. Mesaros, âStrong labeling of sound events using crowdsourced weak labels and annotator competence estimation,â IEEE/ACM Trans. Audio, Speech, Lang. Process., vol. 31, p. 902â914, 2023. [11]Q. Kong, Y. Xu, W. Wang, and M. D. Plumbley, âAudio set classification with attention model: A probabilistic perspective,â in Proc. ICASSP, 2018, p. 316â320. [12]Y. Xu, Q. Kong, W. Wang, and M. D. Plumbley, âLarge-scale weakly supervised audio classification using gated convolutional neural network,â in Proc. ICASSP, 2018, p. 121â125. [13]F. Schmid, P. Primus, T. Morocutti, J. Greif, and G. Widmer, âMulti- iteration multi-stage fine-tuning of transformers for sound event detection with heterogeneous datasets,â in Proc. DCASE, 2024, p. 141â145. [14]H. Nam, D. Min, S. Choi, I. Choi, and Y.-H. Park, âSelf training and ensembling frequency dependent networks with coarse prediction pooling and sound event bounding boxes,â in Proc. DCASE, 2024, p. 96â100. [15]S. Cornell, J. Ebbers, C. Douwes, I. Mart Ì Ä±n-Morat Ì o, M. Harju, A. Mesaros, and R. Serizel, âDcase 2024 task 4: Sound event detection with heterogeneous data and missing labels,â in Proc. DCASE, 2024, p. 31â35. [16] F. Schmid, T. Morocutti, F. Foscarin, J. Schl Ì uter, P. Primus, and G. Widmer, âEffective pre-training of audio transformers for sound event detection,â in Proc. ICASSP, 2025, p. 1â5. [17]P. Cai, Y. Song, K. Li, H. Song, and I. McLoughlin, âMAT-SED: A masked audio transformer with masked-reconstruction based pre-training for sound event detection,â in Proc. Interspeech, 2024. [18]P. Cai, Y. Song, N. Jiang, Q. Gu, and I. McLoughlin, âPrototype based masked audio model for self-supervised learning of sound event detection,â in Proc. ICASSP, 2025, p. 1â5. [19]H. Nam and Y. Park, âJitter: Jigsaw temporal transformer for event reconstruction for self-supervised sound event detection,â CoRR, vol. abs/2502.20857, 2025. [20] X. Li, N. Shao, and X. Li, âSelf-supervised audio teacher-student transformer for both clip-level and frame-level tasks,â IEEE/ACM Trans. Audio, Speech, Lang. Process., vol. 32, p. 1336â1351, 2024. [21] M. Tkachenko, M. Malyuk, A. Holmanyuk, and N. Liubimov, âLabel Studio: Data labeling software,â 2020-2025, open source software available from https://github.com/HumanSignal/label-studio. [22]I. Loshchilov and F. Hutter, âDecoupled weight decay regularization,â in Proc. ICLR, 2019. [23] H. Zhang, M. Ciss Ì e, Y. N. Dauphin, and D. Lopez-Paz, âmixup: Beyond empirical risk minimization,â in Proc. ICLR, 2018. [24]J. Ebbers, R. Haeb-Umbach, and R. Serizel, âThreshold independent evaluation of sound event detection scores,â in Proc. ICASSP, 2022, p. 1021â1025. [25]J. Ebbers, F. G. Germain, G. Wichern, and J. L. Roux, âSound event bounding boxes,â in Proc. Interspeech, 2024.