Paper deep dive
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
Linxi Li, Yuncong Yu, Qianwei Guo, Liwei Jin, Yechen Wang, Carsten Maple
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/7/2026, 3:52:01 PM
Summary
This paper introduces SynSFX, a large-scale dataset for non-speech audio deepfake detection containing 43,374 clips of real and synthetic sound effects generated by seven text-to-audio models. It demonstrates that speech-centric detectors (AASIST, RawNet2) fail on sound effects and suffer catastrophic forgetting when fine-tuned exclusively on them. Joint-domain training with speech data mitigates this forgetting, but zero-shot generalization to unseen generators remains a major bottleneck.
Entities (13)
Relation Signals (13)
SynSFX â contains â 43374 audio clips
confidence 98% ¡ SynSFX comprises a total of 43,374 audio clips, accumulating approximately 178 hours of high-fidelity recordings.
Fine-tuning on SynSFX â causes â Catastrophic Forgetting
confidence 96% ¡ adapting speech-centric detectors exclusively to sound effects causes a catastrophic degradation in speech spoofing detection.
AASIST â evaluatedon â SynSFX
confidence 95% ¡ we evaluated three distinct, representative deepfake detection architectures under a strict zero-shot protocol... AASIST and RawNet2
RawNet2 â evaluatedon â SynSFX
confidence 95% ¡ we evaluated three distinct, representative deepfake detection architectures under a strict zero-shot protocol... AASIST and RawNet2
EAT-AASIST â evaluatedon â SynSFX
confidence 95% ¡ EAT-AASIST shows better zero-shot transfer on SynSFX, achieving an EER of 23.71%
SynSFX â generatedby â Make-An-Audio
confidence 95% ¡ Make-An-Audio [18]
SynSFX â generatedby â TangoFlux
confidence 95% ¡ TangoFlux [20]
SynSFX â generatedby â AudioCraft
confidence 95% ¡ The fake counterparts were generated using seven state-of-the-art text-to-audio (TTA) models: AudioCraft/AudioGen [13]
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects. Existing environmental audio datasets such as EnvSDD provide important initial resources, but remain limited in scale and generation provenance for studying isolated sound-effect deepfakes. To support this direction, we present SynSFX, a large-scale corpus of 43374 clips (26452 synthetic, 16922 real) spanning 7 popular text-to-audio models.
Tags
Links
- Source: https://arxiv.org/abs/2607.04848v1
- Canonical: https://arxiv.org/abs/2607.04848v1
Trouble viewing inline? Open PDF directly â
Full Text
38,945 characters extracted from source content.
Expand or collapse full text
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation Linxi Li ââ * , Yuncong Yu â * , Qianwei Guo â , Liwei Jin â , Yechen Wang â , Carsten Maple â â University of Warwick, Coventry, United Kingdom â OfSpectrum, Inc., Los Angeles, CA, USA AbstractâWhile audio deepfake detection has advanced sig- nificantly, representative detectors show limited generalization to synthetic sound effects. Existing environmental audio datasets such as EnvSDD provide important initial resources, but remain limited in scale and generation provenance for studying isolated sound-effect deepfakes. To support this direction, we present SynSFX, a large-scale corpus of 43374 clips (26452 synthetic, 16922 real) spanning 7 popular text-to-audio models. Index Termsâaudio deepfake detection, sound effects, spoof- ing, non-speech audio, dataset I. INTRODUCTION Rapid advancements in deep generative modeling for audioâencompassing both speech and sound effectsâhave significantly outpaced the development of robust deepfake detection countermeasures [1]. State-of-the-art text-to-audio frameworks now synthesize high-fidelity acoustic events that are frequently perceptually indistinguishable from genuine recordings [2]. However, contemporary deepfake detection literature remains disproportionately anchored to speech syn- thesis and voice cloning paradigms [3]. Consequently, a critical research gap persists regarding non-speech audio forensics, specifically encompassing environmental sounds, Foley, and ambient textures. Despite being equally susceptible to mali- cious manipulation as artificial speech, the targeted detection of synthetic sound effects remains a fundamentally underex- plored domain. However, sound effects differ fundamentally from speech [4]. Unlike speech, which contains lexical and prosodic cues that detectors can exploit, non-speech audio such as sound effects and environmental sounds lacks structured linguistic information. Detection therefore relies purely on acoustic characteristics which modern generative models can closely imitate. As a result, systems trained primarily on speech data often fail to generalize to non-speech scenarios. This limitation is consequential. Sound effects are widely used in gaming, media production, live streaming, and safety- critical systems [5]. As generative models improve, the risk of misuse increasesâfrom subtle manipulation in entertainment to injection of deceptive environmental sounds. To address this gap, we introduce SynSFX (Synthetic Sound Effects), a large- scale corpus for non-speech audio deepfake detection. The dataset includes diverse sound effects and environmental audio * Equal contribution. generated by multiple modern models, establishing a compre- hensive benchmark beyond speech-focused evaluations. The primary contributions of this work are summarized as follows: ⢠A Large-Scale Benchmark with Diagnostic Control: We release SynSFX, a comprehensive 178-hour corpus encompassing seven diverse text-to-audio architectures. Beyond providing massive scale and acoustic diver- sity, SynSFX uniquely features a specialized Shared Prompt Subset. The shared-prompt subset provides a controlled resource for future prompt-matched analyses of generator-dependent artifacts. ⢠Identification of Catastrophic Forgetting: We empiri- cally reveal that adapting speech-centric detectors exclu- sively to sound effects causes a catastrophic degradation in speech spoofing detection. We demonstrate that while joint-domain training mitigates this forgetting, critical generalization bottlenecks remain. ⢠Exposing the Illusion of Generalization: Through rigor- ous zero-shot evaluations on unseen generation architec- tures and feature space visualizations (t-SNE), we show that current models overfit to known synthesis artifacts rather than learning universal acoustic anomalies, estab- lishing a baseline for future generalized audio forensics. I. RELATED WORK A. Speech Deepfake Detection Current research on audio deepfake detection has pre- dominantly focused on human speech. Traditional and learning-based models, such as AASIST and RawNet2, have achieved remarkable success on standardized benchmarks like ASVspoof [6], [7] and FakeAVCeleb [8]. Recent variants, including TO-RawNet [9] and twice-attention networks [10], have further driven Equal Error Rates (EER) down to the 1- 2% range. However, as parallel datasets like WaveFake [11] and JMAD [12] reveal, speech-focused detectors often suffer from a persistent generalization gap under cross-dataset and cross-model transfer, indicating that robust detection remains challenging outside narrow, in-domain conditions. B. Advancements in Audio Generation Models The landscape of synthetic audio has recently expanded far beyond speech. Modern text-to-audio (TTA) and text-to- sound modelsâsuch as AudioCraft [13], AudioLDM variants arXiv:2607.04848v1 [cs.SD] 6 Jul 2026 [14], [15], StableAudio [16], DiffSound [17], Make-An-Audio [18], MMAudio [19], and TangoFlux [20]âcan now generate highly realistic ambient recordings, Foley, and general sound effects directly from text prompts. These modern generative models produce diverse acoustic patterns with significantly fewer spectral artifacts than earlier vocoder-based systems, rendering traditional artifact-driven detection methods increas- ingly obsolete and making deepfake detection exceptionally challenging. C. Deepfake Detection for General Audio and Sound Effects Despite the rapid evolution of general audio generation, the development of countermeasures for non-speech audio is bottlenecked by the lack of dedicated, structurally controlled datasets. Initial efforts, such as the Environmental Sound Deepfake Dataset (EnvSDD) [21], pioneered the concept of environmental sound spoofing by providing a large-scale cor- pus of synthetic soundscapes. Concurrently, the recent Comp- SpoofV2 dataset [22] and the associated ESDD2 challenge [23] significantly advanced the field by addressing component- level spoofingâspecifically, the complex scenario where syn- thetic environmental backgrounds are mixed with authentic foreground human speech. Consequently, state-of-the-art so- lutions like the EAT-AASIST framework [24] have achieved remarkable success in ESDD2 by focusing on separating and classifying these mixed components. However, these foundational works leave a critical method- ological gap in understanding the intrinsic artifacts of gen- erative models. While CompSpoofV2 excels at evaluating mixed-source scenarios, its primary focus is on the artifacts introduced by the composition process, rather than the isolated generative flaws of the TTA models themselves. Conversely, while EnvSDD provides a relatively large volume of isolated synthetic audio, it relies on unconstrained, randomized text prompts across different architectures. This makes prompt- matched cross-generator analysis less direct. To address this gap, our work introduces SynSFX, a large- scale dataset focused on isolated deepfake sound effects. Comprising approximately 178 hours of diverse acoustic envi- ronments, SynSFX includes synthetic audio from seven TTA models with transparent generation provenance. In addition to generator-specific prompts, SynSFX contains a Shared Prompt Subset, where the same text prompts are used across all generators. This subset enables prompt-matched comparisons of detector responses, helping distinguish semantic effects from generator-dependent variation. By focusing on isolated sound effects rather than mixed speech-background scenarios, SynSFX provides a benchmark for studying cross-generator robustness in non-speech audio forensics. I. THE SYNSFX DATASET SynSFX comprises a total of 43,374 audio clips, accumulat- ing approximately 178 hours of high-fidelity recordings. The corpus is strategically designed to encompass both diverse authentic acoustic environments and state-of-the-art synthetic generations under standardized protocols. To systematically evaluate cross-model generalization, SynSFX incorporates synthetic outputs from seven prominent generative architectures. A. Data Sources and Generative Models To ensure a rigorous and balanced evaluation, the dataset is composed of two primary partitions: ⢠Authentic Audio Subset: To serve as the ground truth for deepfake detection, we curated 16,922 real audio clips from five established open-source repositories. The Au- dioCaps and Clotho [25] subsets provide natural environ- mental recordings paired with descriptive captions. ESC- 50 [25] contributes labeled environmental sounds across 50 everyday categories. TACoS [26] adds activity-related audio events, while WavCaps [27] extends coverage to large-scale web-sourced ambient sounds. ⢠SyntheticAudioSubset: The fake counterparts were generated using seven state-of-the-art text-to- audio (TTA) models: AudioLDM (v1/v2) [14], Au- dioCraft/AudioGen [13], MMAudio [19], StableAu- dio [16], Make-An-Audio [18], and TangoFlux [20]. These generators were strategically selected for their architectural diversity, encompassing diffusion-based, transformer-based, and latent generative approaches. In our experimental evaluations, these models are anonymized and denoted as A1âA7 (Table I), respec- tively. B. Prompt Engineering and Expansion Pipeline To ensure immense acoustic diversity and mirror real- world user behaviors, SynSFX utilizes a structured prompt generation pipeline driven by advanced Large Language Mod- els (LLMs), including ChatGPT and Gemini. The workflow initiates from concise baseline descriptions (e.g., âfootsteps on gravelâ), which are subsequently expanded into contextually rich textual prompts (e.g., âa person walking briskly on a gravel path under light rainâ). This expansion mechanism introduces granular environmental variables and scene com- plexities, thereby enhancing the perceptual realism of the syn- thesized events. All expanded prompts undergo strict human filtration to eliminate semantic ambiguities or unsafe content. C. Corpus Architecture and Technical Specifications The prompt distribution within SynSFX utilizes a total of 28,350 unique textual prompts, mathematically structured to facilitate both comprehensive training and rigorous cross- model comparative analysis: ⢠Shared Prompt Subset: Comprises 1,890 identical prompts provided universally to all evaluated TTA mod- els. This subset serves as a controlled baseline, allowing researchers to isolate semantic variables and perform direct, cross-architecture comparisons of generation ar- tifacts. ⢠Exclusive Prompt Subsets: The remaining distribution is evenly partitioned, with each generative model producing audio clips from unique prompts (detailed in Table I) exclusive to that specific architecture, guaranteeing sta- tistical balance. To preserve inherent generation artifacts, all audio clips are stored in uncompressed WAV format using the native sample rate of their source. Diffusion-based generators (AudioCraft, AudioLDM variants, Make-An-Audio) produce 16 kHz audio, while M-Audio, StableAudio, and TangoFlux output at 44.1 kHz. The real subsets are primarily recorded at 44.1 kHz, with the exception of AudioCaps at 22.0 kHz. For benchmarking, SynSFX is released with predefined train, validation, and test splits to ensure standardized evaluation across future research. IV. EXPERIMENTS & RESULTS A. Baselines To establish a benchmark on SynSFX, we evaluated three distinct, representative deepfake detection architectures un- der a strict zero-shot protocol (i.e., evaluated directly on SynSFX without any domain-specific fine-tuning). The se- lected baselines include two well-established speech-centric models (AASIST and RawNet2) and a recent state-of-the- art framework capable of generalized audio processing (EAT- AASIST). The evaluation results are detailed in Table I. B. Collapse of Speech-Centric Detectors As summarized in Table I, both AASIST and RawNet2 ex- hibit severe performance degradation, performing near random chance levels: RawNet2 yields an Equal Error Rate (EER) of 49.94%, while AASIST collapses entirely with an EER of 60.84% and an exceptionally high False Acceptance Rate (FAR). This profound confusion pattern is mechanically pre- dictable. Both AASIST and RawNet2 were originally op- timized for the ASVspoof benchmarks, heavily relying on speech-intrinsic properties such as fundamental frequency subbands, vocal tract formants, and linguistic prosody. Be- cause the SynSFX corpus fundamentally lacks these structured lexical and human-vocal cues, the traditional detectors fail to anchor onto any meaningful features, resulting in random feature-space distributions. C. Partial Success and Persistent Gaps of Generalized Model While pure speech-trained models perform poorly, EAT- AASIST shows better zero-shot transfer on SynSFX, achiev- ing an EER of 23.71% and an ROC AUC of 0.8590. This relative robustness is likely due to its EAT backbone and prior training on ESDD2 [23], where the model is exposed to synthetic environmental components mixed with speech. However, its performance remains below the in domain re- sults, indicating that mixed-audio environmental spoofing and isolated sound-effect deepfake detection are not equivalent. These results motivate SynSFX as a dedicated benchmark for studying non-speech audio deepfake detection under isolated and multi-generator conditions. TABLE I MODELS AND DATASETS USED IN CONSTRUCTING THE SYNSFX CORPUS. DURATIONS ARE IN HOURS. THE SHARED PROMPT SUBSET CONTAINS 1,890 PROMPTS THAT ARE GENERATED BY EACH AI GENERATOR; THESE CLIPS ARE INCLUDED IN EACH MODELâS TOTAL AND ARE NOT COUNTED AS AN ADDITIONAL GENERATOR. Model / DatasetSharedModel-specific# ClipsDuration [h] A1 AudioCraft1890189037807.9 A2 AudioLDM11890189037807.8 A3 AudioLDM21890189037807.9 A4 MMAudio1890189037806.8 A5 Make-An-Audio18901890378011.3 A6 Stable Audio18901888377810.5 A7 TangoFlux1890188437748.9 Total Generated13230132222645261.1 AudioCapsâ400011.0 Clothoâ383924.0 ESC-50â20002.8 TACoSâ500031.1 WavCapsâ208350.9 Total Realâ16922119.7 TABLE I ZERO-SHOT EVALUATION OF BASELINE DEEPFAKE DETECTORS ON SYNSFX EVALUATION SUBSET(WITHOUT FINE-TUNING) MetricAASISTRawNet2EAT-AASIST True Fake (TP)369847277205 False Real (FN)574647172239 False Fake (FP)1029584514012 True Real (TN)6627847112910 Equal Error Rate (EER)60.84%49.94%23.71% Threshold at EER0.00210-0.021420.11014 ROC AUC0.36090.50340.8498 PR AUC0.27440.39020.7902 F1 Score0.31580.41790.6974 False Acceptance Rate (FAR)0.60820.49940.2371 False Rejection Rate (FRR)0.60820.49950.2371 D. Finetuning Protocol To investigate whether domain-specific training can over- come the limitations observed in the zero-shot evaluations, we fine-tuned both AASIST and EAT-AASIST on the SynSFX corpus. Both architectures were optimized with binary su- pervision, classifying inputs as authentic or synthetic. To prevent data leakage, all clips were assigned globally unique identifiers, ensuring mutually exclusive training, validation, and evaluation splits. To reduce preprocessing-related shortcuts, all audio was converted to mono, resampled to 16 kHz, and peak-normalized before training and evaluation. AASIST and EAT-AASIST used fixed input lengths of 64,600 and 64,000 samples, respec- tively; longer clips were cropped (random crop during training and center crop during validation/testing), while shorter clips were symmetrically zero-padded. Both models were initialized from their official pre-trained weights [24], [28] and fine-tuned end-to-end for 60 epochs on a single RTX 4090D GPU without architectural modification. We used AdamW with a learning rate of 1Ă 10 â4 , weight decay of 1Ă 10 â5 , and CosineAnnealingWarmRestarts (T 0 = 10,T mult = 2). Batch sizes were 32 for AASIST and 16 for EAT-AASIST. E. Evaluation To systematically quantify both in-domain retention and cross-domain generalization, we established a rigorously strat- ified evaluation protocol consisting of two mutually exclusive test sets. In-Domain (Seen Generators) Test Set: To measure base- line fine-tuning efficacy, we extracted a dedicated test split directly from the SynSFX corpus. This in-domain evaluation set comprises 4,338 audio clips, totaling approximately 16 hours of acoustic data. While strictly excluded from training dataset at sample level, this subset exposes the detectors to the same acoustic domains, prompt structures, and text- to-audio architectures (A1âA7) encountered during the fine- tuning phase. Out-of-Domain (Unseen Generators) Test Set: To rigor- ously assess cross-model transferability and zero-shot robust- ness, we constructed an entirely independent out-of-domain evaluation set comprising 1,113 audio clips. To guarantee a statistically sound evaluation, this set is symmetrically parti- tioned to maintain a strict 1:1 class balance: ⢠Synthetic Subset (50%): Generated using a state-of-the- art, proprietary commercial text-to-audio API. By utiliz- ing a closed-source architecture entirely absent from the SynSFX training phase, we ensure the detection models have zero prior exposure to its specific synthesis artifacts. ⢠Authentic Subset (50%): Sampled exclusively from the UrbanSound8K [29] dataset. This provides a distinct distribution of real-world environmental recordings and Foley sound effects that remain completely unrepresented in the primary training corpus. By evaluating the fine-tuned models across these two dis- tinct sets, we can accurately decouple genuine acoustic repre- sentation learning from mere overfitting to generator-specific flaws. F. In-Domain Efficacy and Catastrophic Forgetting As detailed in Table I, fine-tuning AASIST exclusively on the SynSFX corpus yields substantial performance im- provements within the non-speech domain. The standard A- SIST model successfully reduces its Equal Error Rate (EER) from the near-random 60.84% (zero-shot) to an impressive 3.23%. Similarly, the more advanced EAT-AASIST framework achieves a highly discriminative EER of 2.36% with an F1 score of 0.9806. These in-domain results validate the utility of the SynSFX dataset, demonstrating that synthetic sound effects possess learnable, domain-specific generation artifacts that deep architectures can successfully extract when provided with sufficient domain-targeted supervision. However, this dramatic improvement in sound effect foren- sics comes at a severe cost to the modelsâ original capabilities. When the SynSFX-only fine-tuned AASIST model is evalu- ated back on its original speech deepfake testing set [6], we observe a phenomenon of catastrophic forgetting. The speech evaluation EER spikes to 33.61% (Table I), indicating a profound degradation in voice spoofing detection. This functional collapse mechanically suggests that the network shifts its decision boundary to optimize for the heterogeneous spectral noise, diverse sound-effect spectra, and generator-specific flaws in synthetic effects, thereby âunlearn- ingâ the more consistent harmonic structures and vocal-tract cues required for authenticating human speech. This trade- off confirms that adapting a single-domain model via naive fine-tuning is inherently insufficient for generalized audio forensics, necessitating joint-domain training paradigms. G. Mitigating Forgetting via Joint-Domain Training To address the severe degradation in speech forensics, we investigated a joint-domain training paradigm. The objective is to determine whether foundational acoustic models can maintain dual decision boundariesâpreserving human vocal tract priors while simultaneously learning the heterogeneous artifacts of synthetic environmental sounds. To facilitate this, we integrated a large-scale speech deep- fake corpus derived from the ASVspoof 2019 Logical Access (LA) dataset [6]. The joint fine-tuning partition incorporates 25,380 speech utterances (24.15 hours, comprising 22,800 synthetic and 2,580 authentic clips), while validation utilizes 24,844 utterances (24.00 hours, 22,296 synthetic and 2,548 authentic). For a rigorous speech-domain evaluation, we uti- lized an isolated evaluation split comprising 71,237 utterances (61.50 hours). This massive test set ensures the models are evaluated against diverse, unseen speech spoofing algorithms that were strictly excluded from the training phase. As demonstrated in Table I, interleaving speech data into the fine-tuning process successfully rescues both architectures from catastrophic forgetting. For the standard AASIST model, joint-domain training restores the speech evaluation EER to an exceptional 3.61% (a dramatic recovery from 33.61%), while simultaneously maintaining a highly competitive EER of 3.76% on the SynSFX test set. The EAT-AASIST framework exhibits an even stronger capacity for multi-domain representation. Under the joint- training paradigm, it achieves a 2.45% EER and an F1 score of 0.9798 on synthetic sound effects, alongside a restored 5.25% EER on speech. These results empirically confirm that joint-domain training acts as an effective stabilizer. It suggests that the network has the capacity to simultaneously map the acoustic signatures of both domains without mutual exclusion, even with minor domain-specific performance degradation. However, while this joint paradigm successfully solves the issue of cross-domain memory retention, it introduces the final and most critical challenge of audio deepfake detection: zero- shot generalization to entirely unseen generators. H. The Illusion of Generalization: Overfitting to Generator Artifacts While joint-domain training mitigates cross-domain forget- ting, Table I shows that zero-shot generalization to unseen generators remains the dominant bottleneck. To separate fake- generator shift from real-domain shift, we evaluate two diag- nostic settings: Seen Real + Unseen Fake and Unseen Real + TABLE I COMPREHENSIVE CROSS-DOMAIN EVALUATION OF FINE-TUNED DETECTORS. SynSFX-Only USES ONLY SOUND-EFFECT DATA, WHILE SynSFX + Speech USES JOINT-DOMAIN TRAINING. THE UNSEEN EVALUATION IS FURTHER DECOMPOSED TO SEPARATE FAKE-GENERATOR SHIFT FROM REAL-DOMAIN SHIFT. ArchitectureTraining ParadigmEvaluation SetEERROC AUCPR AUCF1 Score AASIST SynSFX-Only SynSFX Test (Seen)3.23%0.99590.99750.9733 Speech (Out-of-Domain)33.61%0.73590.96260.7799 Seen Real + Unseen Fake30.21%0.76390.55290.5326 Unseen Real + Seen Fake4.28%0.99550.99900.9738 Unseen Real + Unseen Fake26.17%0.79600.78170.7381 SynSFX + Speech SynSFX Test (Seen)3.76%0.99270.99410.9689 Speech (In-Domain)3.61%0.99420.99930.9795 Seen Real + Unseen Fake34.33%0.72280.46750.4854 Unseen Real + Seen Fake5.43%0.98780.99740.9663 Unseen Real + Unseen Fake37.23%0.66870.63770.6277 EAT-AASIST Pre-trained BaselineSpeech (Source Domain)4.05%0.99320.99920.9770 SynSFX-Only SynSFX Test (Seen)2.36%0.99780.99860.9806 Speech (Out-of-Domain)24.04%0.84250.97970.8500 Seen Real + Unseen Fake25.63%0.82220.60060.5887 Unseen Real + Seen Fake2.16%0.99760.99950.9868 Unseen Real + Unseen Fake30.13%0.76600.76840.6990 SynSFX + Speech SynSFX Test (Seen)2.45%0.99780.99860.9798 Speech (In-Domain)5.25%0.98910.99870.9700 Seen Real + Unseen Fake27.50%0.79890.58310.5652 Unseen Real + Seen Fake2.16%0.99760.99950.9868 Unseen Real + Unseen Fake32.19%0.74970.75780.6781 Seen Fake. Across both architectures and training paradigms, performance degrades substantially when the fake samples come from the unseen generator. For example, joint-trained AASIST reaches 34.33% EER on Seen Real + Unseen Fake, while joint-trained EAT-AASIST obtains 27.50% EER. In contrast, replacing the real side with UrbanSound8K while keeping seen SynSFX generators remains much less harmful: AASIST and EAT-AASIST achieve 5.43% and 2.16% EER, respectively, under Unseen Real + Seen Fake. This decomposition indicates that the main source of out-of- domain failure is the unseen synthetic generator rather than the authentic-source shift alone. When both shifts are combined in Unseen Real + Unseen Fake, performance further de- grades to 37.23% EER for joint-trained AASIST and 32.19% EER for joint-trained EAT-AASIST. These results suggest that current detectors still rely heavily on generator-specific artifacts learned from the seen SynSFX models, limiting their robustness to novel TTA architectures. Feature Space Visualization: To visually substantiate the challenge of general sound effect forensics, we projected the penultimate-layer embeddings of both joint-trained models into a 2D plane using t-SNE (Figure 1). The projection contrasts four subsets: authentic speech, synthetic speech, real sound effects (UrbanSound8K), and unseen synthetic sound effects (A8). The visualizations highlight a stark contrast between the speech and non-speech domains. While AASIST and EAT- AASIST exhibit different mapping behaviorsâforming dis- crete clusters and a continuous U-shaped manifold, respec- tivelyâboth models successfully establish clear linear sep- arability between authentic and synthetic speech. Crucially, this discriminative structure entirely collapses for unseen non- speech audio. In both plots, real sound effects (green) and unseen synthetic effects (purple) are catastrophically entangled with no discernible decision boundary. This performance degradation indicates that current deep- fake detection paradigms suffer from an âillusion of general- ization.â Mechanically, the models are not learning the funda- mental, physics-based acoustic discrepancies between natural and synthesized environments. Instead, they are overfitting to the specific digital signatures, phase distortions, and diffusion noises inherent to the seven generative architectures (A1â A7) encountered during training. When confronted with a novel commercial architecture lacking these exact spectral fingerprints, the decision boundaries transfer poorly. Aligning with Broader Forensics Challenges: This phe- nomenon is not unique to non-speech audio; it perfectly mirrors the well-documented generalization gaps historically observed in speech anti-spoofing. Foundational studies eval- uating cross-corpus robustness (such as the ASVspoof 2021 Challenge [7] and cross-database evaluations on WaveFake [11]) have repeatedly demonstrated that models optimized on specific neural vocoders frequently degrade to near-random chance when tested on unseen generation algorithms. Our â40â200204060 t-SNE 1 â40 â20 0 20 40 60 t-SNE 2 AASIST (SynSFX+Speech fine-tuned) real_speech (n=300) asvspoof_spoof (n=300) unseen generator (n=300) real_sfx_urban8k (n=300) (a) AASIST â60â40â20020406080 t-SNE 1 â30 â20 â10 0 10 20 30 40 t-SNE 2 EAT-AASIST (SynSFX+Speech fine-tuned) real_speech (n=300) asvspoof_spoof (n=300) unseen generator (n=300) real_sfx_urban8k (n=300) (b) EAT-AASIST Fig. 1. t-SNE visualization of the penultimate-layer embeddings for the joint-trained (a) AASIST and (b) EAT-AASIST models. Both architectures maintain linear separability for speech domains (red vs. blue) but suffer from severe feature entanglement when processing unseen non-speech audio (green vs. purple). empirical results confirm that this âoverfitting to generator arti- factsâ is equally, if not more, severe in the highly unstructured domain of general sound effects. Shared-Prompt Diagnostic Analysis Since EAT-AASIST already shows partial zero-shot discrimination on SynSFX before fine-tuning (Section IV-C), we use it to analyze the shared-prompt subset. For each shared prompt, we compute the variance and range of fake scores across the seven gen- erators, and then average these prompt-level statistics over all shared prompts. Before adaptation, EAT-AASIST obtains an overall fake-score mean of 0.5135, its per-prompt range remains large (mean/median: 0.8915/0.9461), indicating that detector confidence varies strongly across generators even under identical semantic prompts. After SynSFX fine-tuning, the overall fake-score mean increases to 0.9865, while the per-prompt variance mean drops from 0.1413 to 0.0082 and the range mean/median decreases to 0.0899/0.0037. This suggests that SynSFX supervision sub- stantially improves cross-generator consistency for seen gener- ators, although residual generator-specific or prompt-generator interaction effects may still influence detector confidence. The Foundational Value of SynSFX: The cross-model collapse in Section IV-H illuminates the critical necessity of the SynSFX corpus. It suggests that naive fine-tuning isnât sufficient for robust audio forensics. To break this bottleneck, the community must transition toward learning generalized synthetic acoustic representations. SynSFX is strategically engineered to facilitate this exact paradigm shift through its Shared Prompt Subset. While different text-to-audio models interpret and render semantics with inherent stochastic variance, this shared subset provides an mechanism to control for macro-semantic variables. By ensuring that identical acoustic events (e.g., a âfootstepâ) are represented across diverse generator distributions, it actively TABLE IV SHARED-PROMPT DIAGNOSTIC ANALYSIS WITH EAT-AASIST. ModelOverall fake scoreVar. meanRange mean / median Pre-trained0.51350.14130.8915 / 0.9461 SynSFX fine-tuned0.98650.00820.0899 / 0.0037 mitigates content-bias. This design encourages future methods to evaluate detector behavior under matched semantic content and to better separate generator-dependent cues from acoustic- class effects. V. CONCLUSION In this work, we introduced SynSFX, a multi-generator benchmark for isolated non-speech audio deepfake detection. Our experiments show that joint-domain training can mitigate speech-domain forgetting while maintaining strong perfor- mance on seen SynSFX generators. However, performance still degrades under unseen-generator evaluation, suggesting that current detectors remain sensitive to generator- and source- specific artifacts. By releasing SynSFX with transparent provenance and a shared-prompt subset, we aim to support controlled studies of semantic content, generator effects, and cross-domain ro- bustness. We also acknowledge several limitations: the dataset cannot cover the full diversity of real-world soundscapes, the unseen-generator setting remains limited in scale, and resolving cross-generator generalization is beyond the scope of this work. Future work will extend the benchmark and explore more robust representation learning methods. VI. DATA AVAILABILITY The full dataset is available at https://ofspectrum.com/news/ synsfx. VII. GENERATIVE AI USE DISCLOSURE Large Language Models (LLMs) were used solely for manuscript polishing (e.g., rephrasing and grammar checks) to improve clarity and readability. The LLMs were not used for ideation, methodology, experimental design, data analysis, or result interpretation. All scientific content was produced and verified by the authors. REFERENCES [1] Z. Khanjani, G. Watson, and V. P. Janeja, âAudio deepfakes: A survey,â Frontiers in Big Data, vol. Volume 5 - 2022, 2023. [Online]. Available: https://w.frontiersin.org/journals/big- data/articles/10.3389/fdata.2022.1001063 [2] A. Vyas, B. Shi, M. Le, A. Tjandra, Y.-C. Wu, B. Guo, J. Zhang, X. Zhang, R. Adkins, W. Ngan, J. Wang, I. Cruz, B. Akula, A. Akinyemi, B. Ellis, R. Moritz, Y. Yungster, A. Rakotoarison, L. Tan, C. Summers, C. Wood, J. Lane, M. Williamson, and W.-N. Hsu, âAudiobox: Unified audio generation with natural language prompts,â 2023. [Online]. Available: https://arxiv.org/abs/2312.15821 [3] X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. Kinnunen, M. Todisco, J. Yamagishi, N. Evans, A. Nautsch, and K. A. Lee, âAsvspoof 2021: Towards spoofed and deepfake speech detection in the wild,â IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, p. 2507â2522, 2023. [Online]. Available: http://dx.doi.org/10.1109/TASLP.2023.3285283 [4] S. Chachada and C.-C. J. Kuo, âEnvironmental sound recognition: A survey,â in 2013 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, 2013, p. 1â9. [5] M. Crocco, M. Cristani, A. Trucco, and V. Murino, âAudio surveillance: a systematic review,â 2014. [Online]. Available: https: //arxiv.org/abs/1409.7787 [6] M. Todisco, X. Wang, V. Vestman, M. Sahidullah, H. Delgado, A. Nautsch, J. Yamagishi, N. W. D. Evans, T. H. Kinnunen, and K. A. Lee, âASVspoof 2019: Future horizons in spoofed and fake audio detection,â in Proceedings of INTERSPEECH 2019, 2019, p. 1008â1012. [Online]. Available: https://dblp.org/rec/conf/interspeech/ Todisco0VSDNYEK19.html [7] H. Delgado, N. W. D. Evans, T. Kinnunen, K. A. Lee, X. Liu, A. Nautsch, J. Patino, M. Sahidullah, M. Todisco, X. Wang, and J. Yamagishi, âASVspoof 2021: Automatic speaker verification spoofing and countermeasures challenge evaluation plan,â ASVspoof Consortium, Evaluation Plan / Technical Report, Jul. 2021. [8] H. Khalid, S. Tariq, M. Kim, and S. S. Woo, âFakeavceleb: A novel audio-video multimodal deepfake dataset,â 2022. [Online]. Available: https://arxiv.org/abs/2108.05080 [9] C. Wang, J. Yi, J. Tao, C. Zhang, S. Zhang, R. Fu, and X. Chen, âTo-rawnet: Improving RawNet with TCN and orthogonal regularization for fake audio detection,â 2023. [Online]. Available: https://arxiv.org/abs/2305.13701 [10] D. Yao, X. Yuan, X. Hu, and G. Guo, âTwice attention networks for synthetic speech detection,â Neurocomputing, vol. 559, p. 126799, 2023. [Online]. Available: https://doi.org/10.1016/j.neucom.2023.126799 [11] J.FrankandL.Sch Ě onherr,âWavefake:Adatasetto facilitateaudiodeepfakedetection,âinProceedingsofthe Neural Information Processing Systems Track on Datasets and Benchmarks 1 (NeurIPS Datasets and Benchmarks 2021), 2021. [Online]. Available: https://datasets-benchmarks-proceedings.neurips.c/ paper/2021/file/c74d97b01eae257e44a9d5bade97baf-Paper-round2.pdf [12] C. O. Mawalim, Y. Wang, A. Adila, S. Okada, and M. Unoki, âMulti- lingual deepfake speech dataset for robust and generalizable detection,â IEEE Access, vol. 14, p. 57 144â57 161, 2026. [13] J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. D Ě efossez, âSimple and controllable music generation,â in Proceedings of the 37th International Conference on Neural Information Processing Systems, ser. NIPS â23. Red Hook, NY, USA: Curran Associates Inc., 2023. [14] H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. P. Mandic, W. Wang, and M. D. Plumbley, âAudioldm: Text-to-audio generation with latent diffusion models,â 2023. [Online]. Available: https: //arxiv.org/abs/2301.12503 [15] H. Liu, Y. Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y. Wang, W. Wang, Y. Wang, and M. D. Plumbley, âAudioldm 2: Learning holistic audio generation with self-supervised pretraining,â 2023. [Online]. Available: https://arxiv.org/abs/2308.05734 [16] Stability AI, âStable audio: Fast timing-conditioned latent audio diffusion,â Stability AI, Technical Report (web publication), Sep. 2023. [Online]. Available: https://stability.ai/research/stable-audio-fast- timing-conditioned-latent-audio-diffusion [17] D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y. Zou, and D. Yu, âDiffsound: Discrete diffusion model for text-to-sound generation,â IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, p. 1720â1733, 2023. [Online]. Available: https://dl.acm.org/doi/10.1109/TASLP.2023.3268730 [18] R. Huang, J. Huang, D. Yang, Y. Ren, L. Liu, M. Li, Z. Ye, J. Liu, X. Yin, and Z. Zhao, âMake-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,â in Proceedings of the 40th International Conference on Machine Learning (ICML), ser. Proceedings of Machine Learning Research, vol. 202, 2023, p. 13 916â13 932. [Online]. Available: https://proceedings.mlr.press/v202/huang23i.html [19] H. K. Cheng, M. Ishii, A. Hayakawa, T. Shibuya, A. Schwing, and Y. Mitsufuji, âMmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis,â 2025. [Online]. Available: https://arxiv.org/abs/2412.15322 [20] C.-Y. Hung, N. Majumder, Z. Kong, A. Mehrish, A. A. Bagherzadeh, C. Li, R. Valle, B. Catanzaro, and S. Poria, âTangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,â 2025. [Online]. Available: https://arxiv.org/abs/2412.21037 [21] H. Yin, Y. Xiao, R. K. Das, J. Bai, H. Liu, W. Wang, and M. D. Plumbley, âEnvsdd: Benchmarking environmental sound deepfake detection,â 2025. [Online]. Available: https://arxiv.org/abs/2505.19203 [22] X. Zhang, Y. Wang, L. Li, L. Jin, and M. Li, âCompspoof: A dataset and joint learning framework for component-level audio anti-spoofing countermeasures,â 2026. [Online]. Available: https: //arxiv.org/abs/2509.15804 [23] X. Zhang, H. Yin, Y. Xiao, L. Zhang, T. Dang, R. K. Das, and M. Li, âOverview of esdd2: Environment-aware speech and sound deepfake detection challenge,â 2026. [Online]. Available: https://arxiv.org/abs/2606.10791 [24] J. Cao, C. Fan, J. Xue, Y. Xie, R. Fu, Z. Wen, J. Yi, Y. Ren, Z. Lv, and J. Tao, âEfficient audio transformer and aasist for environment sound deepfake detection in the esdd 2026 challenge,â in ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, p. 21 781â21 783. [25] K. Drossos, S. Lipping, and T. Virtanen, âClotho: an audio captioning dataset,â in 2020 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2020, Barcelona, Spain, May 4-8, 2020.Barcelona, Spain: IEEE, May 2020, p. 736â740. [Online]. Available: https://doi.org/10.1109/ICASSP40776.2020.9052990 [26] P. Primus, F. Schmid, and G. Widmer, âTACOS: Temporally-aligned audio captions for language-audio pretraining,â in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA 2025, Tahoe City, CA, USA, October 12-15, 2025.Tahoe City, CA, USA: IEEE, October 2025, p. 1â5. [Online]. Available: https://doi.org/10.1109/WASPAA66052.2025.11230997 [27] X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang, âWavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,â IEEE/ACM Transactions on Audio, Speech and Language Processing, vol. 32, p. 3339â3354, 2024. [Online]. Available: https://doi.org/10.1109/TASLP.2024.3419446 [28] J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.-J. Yu, and N. W. D. Evans, âAASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,â in 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, p. 6367â6371. [Online]. Available: https://dblp.org/rec/conf/icassp/JungHTSCLYE22.html [29] J. Salamon, C. Jacoby, and J. P. Bello, âA dataset and taxonomy for urban sound research,â in Proceedings of the 22nd ACM International Conference on Multimedia (M â14), 2014, p. 1041â1044. [Online]. Available: https://dl.acm.org/doi/10.1145/2647868.2655045