Paper deep dive
A Production-Oriented Framework for Evaluation of SFX Generation
Mélodie Desbos, Yara Bahram, Eric Granger, Mohammadhadi Shateri
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/14/2026, 4:28:28 AM
Summary
The paper introduces a production-oriented evaluation framework for reference-guided sound effects (SFX) generation, addressing limitations in existing task-specific assessments. It defines nine production requirements and a two-stage protocol combining a shared audio-to-audio (ATA) variation task on the ESC-50 dataset with capability-specific analyses. The framework evaluates five baselines (AudioLDM, AudioX, T-Foley, ThinkSound, A2SB) using objective metrics (FAD, ImageBind alignment, S-MOS, transient diagnostics) and human studies. Results indicate AudioX offers the best trade-off between reference alignment and diversity, while other models excel in specific editing operations like temporal control, targeted modification, or inpainting, providing a structured decision protocol for industrial audio pipelines.
Entities (11)
Relation Signals (15)
Production-Oriented Evaluation Framework → usesdataset → ESC-50
confidence 98% · Experiments use ESC-50, with 2,000 five-second recordings across 50 classes.
Production-Oriented Evaluation Framework → evaluates → AudioX
confidence 95% · Among the full-generation baselines evaluated under a shared ATA setting, AudioX provides the strongest overall trade-off between reference alignment and diversity...
Production-Oriented Evaluation Framework → evaluates → T-Foley
confidence 95% · T-Foley (5) for time-conditioned control; ThinkSound (19) for targeted editing; and A2SB (12) for inpainted variations.
Production-Oriented Evaluation Framework → evaluates → AudioLDM
confidence 95% · We evaluate five representative baselines covering complementary SFX variation capabilities. AudioLDM (17) and AudioX (28) for SFX morphing...
Production-Oriented Evaluation Framework → evaluates → ThinkSound
confidence 95% · ThinkSound (19) for targeted editing...
Production-Oriented Evaluation Framework → evaluates → A2SB
confidence 95% · A2SB (12) for inpainted variations.
AudioX → providesbesttradeofffor → Reference alignment and diversity
confidence 94% · Among the full-generation baselines evaluated under a shared ATA setting, AudioX provides the strongest overall trade-off between reference alignment and diversity while still supporting SFX morphing.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows. Existing evaluations are usually tied to text-to-audio (TTA), unconditional, or task-specific settings, limiting assessment for reference-guided sound effects (SFX) variation. To address this gap, we present a production-oriented evaluation framework for structured comparison of heterogeneous audio generation and editing methods. Our framework identifies nine production requirements and explicitly accounts for differences in model capabilities, enabling comparison under a common production objective. A two-stage protocol is introduced: (1) a reference-guided audio-to-audio (ATA) variation task, in which all methods are evaluated under the same ESC-50 SFX adaptation setup, and (2) capability-specific analyses of native operations such as SFX morphing, temporal and energy alignment, inpainting, and targeted editing. This framework combines objective metrics (including FAD, ImageBind-based reference alignment, and diversity across generated variants), together with a human study of perceptual identity preservation and transient diagnosis. Our study reveals complementary strengths and trade-offs across baselines for different production needs. Among the full-generation baselines evaluated under a shared ATA setting, AudioX provides the strongest overall trade-off between reference alignment and diversity while still supporting SFX morphing. Other baselines remain most suitable for specific editing operations. Our framework establishes a structured evaluation and decision protocol for reference-guided SFX variation and provides a practical basis for designing future unified industrial audio generation pipelines. Audio demos are on the accompanying web page.
Tags
Links
- Source: https://arxiv.org/abs/2607.09973v1
- Canonical: https://arxiv.org/abs/2607.09973v1
Trouble viewing inline? Open PDF directly →
Full Text
73,798 characters extracted from source content.
Expand or collapse full text
A Production-Oriented Framework for Evaluation of SFX Generation Abstract Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows. Existing evaluations are usually tied to text-to-audio (TTA), unconditional, or task-specific settings, limiting assessment for reference-guided sound effects (SFX) variation. To address this gap, we present a production-oriented evaluation framework for structured comparison of heterogeneous audio generation and editing methods. Our framework identifies nine production requirements and explicitly accounts for differences in model capabilities, enabling comparison under a common production objective. A two-stage protocol is introduced: (1) a reference-guided audio-to-audio (ATA) variation task, in which all methods are evaluated under the same ESC-50 SFX adaptation setup, and (2) capability-specific analyses of native operations such as SFX morphing, temporal and energy alignment, inpainting, and targeted editing. This framework combines objective metrics (including FAD, ImageBind-based reference alignment, and diversity across generated variants), together with a human study of perceptual identity preservation and transient diagnosis. Our study reveals complementary strengths and trade-offs across baselines for different production needs. Among the full-generation baselines evaluated under a shared ATA setting, AudioX provides the strongest overall trade-off between reference alignment and diversity while still supporting SFX morphing. Other baselines remain most suitable for specific editing operations. Our framework establishes a structured evaluation and decision protocol for reference-guided SFX variation and provides a practical basis for designing future unified industrial audio generation pipelines. Audio demos and further details can be found on the accompanying web page111https://melodiedesbos.github.io/sfx-eval-framework. 1 Introduction Production Motivation — In SFX production, sound designers often rely on few reusable recordings, since manually designing variations is costly and time-consuming. This low variability can lead to perceptible repetition in interactive media, motivating systems that can produce useful variations from a reference audio clip (5; 14; 7; 15; 19) or from standard TTA generation (6; 18; 13), rather than synthesizing audio unconditionally (39). In practice, an effective sound design system must do more than produce plausible audio as it is subject to multiple constraints. Unlike generic TTA generation, this setting requires jointly balancing realism, identity preservation, controllability, and usability. Gaps in Current Evaluations — Recent audio generation progress has made SFX generation increasingly accessible. Diffusion, flow-matching, and multimodal pipelines now support style transfer (17; 28) (or SFX morphing (23)), temporal conditioning (5), targeted editing (19), and inpainting (12). However, these methods are typically evaluated only within their original task settings, such as TTA generation (17; 6), video-to-audio (VTA) synthesis (19; 21), foley alignment (5; 3), or localized restoration (12). Therefore, their reported performance offers limited insight into controlled and production-useful reference-guided variation. In addition, differences in task formulation, architecture, and evaluation criteria make methods difficult to compare and complicate the selection of the most suitable approach for specific SFX generation needs. Proposal of Evaluation Framework — This work aims to address this gap by proposing a production-oriented evaluation framework for reference-guided SFX variation. The proposed framework defines practical production requirements, selects complementary state-of-the-art baselines accordingly, and evaluates them through a two-stage protocol: (1) All methods are compared under a shared reference-conditioned ATA variation task using ESC-50, an open-source few-shot soundbank, to reflect a realistic production usage (25). (2) Since the evaluated baselines differ in conditioning, training, and editing capabilities, our goal is not to produce a single ranking but to characterize their suitability for creative SFX variation. The proposed framework further provides a capability-specific analyses to assess strengths and limitations of native editing operations such as reference-guided SFX morphing, temporal alignment, inpainting, and targeted editing. This design enables a structured and fair comparison of heterogeneous methods under a shared production objective, while creating a decision protocol to emphasize distinctive strengths of each approach. Contributions — (i) We introduce a production-oriented evaluation framework for reference-guided SFX variation grounded in practical requirements such as identity preservation, diversity, controllable variation, and production usability. (i) A two-stage protocol is proposed that combines a shared reference-conditioned ATA variation task with capability-specific analyses. This enables a decision-oriented analysis and more structured comparison across heterogeneous baselines. (i) Comparison of recent audio generation and editing methods through objective, perceptual, qualitative, and capability-specific evaluations. It reveals clear trade-offs and complementary strengths of different methods for reference-guided SFX variation. 2 Related Work 2.1 Audio Generation for SFX In audio generation, diffusion-based models remain a dominant paradigm for high-fidelity synthesis and are increasingly used for unconditional or text-conditioned synthesis, while operating on waveform, spectrogram, or latent representations (17; 5; 22). AudioLDM adopts a latent-diffusion formulation in a VAE-based latent space and relies on CLAP embeddings for audio-text conditioning (17). T-Foley (5) is guided by sound-class and temporal-event information for Foley generation. Recently, flow-matching models have also been explored for controllable audio generation, conditioning the flow on multimodal semantic embeddings (19). Another common denoising architecture is the Diffusion Transformer (DiT), in which self-attention operates on multimodal features together with audio tokens or latent audio representations, while the Transformer serves as the denoising network (28; 8; 3; 37; 26; 4). Earlier SFX-oriented works also studied controllable SFX generation, through synthesis of environmental sounds from onomatopoeic words (24), or from acoustic features with explicit control over pitch, loudness, timbre, and transients (20). These works motivate SFX generation as controllable synthesis, but lack inputs over recent generative strengths. 2.2 Reference-Guided Audio Variation and Editing Beyond unconditional or TTA synthesis, audio editing aims to modify existing samples while preserving timbre, temporal structure, or event identity. Recent generative audio frameworks support editing driven either by explicit instructions (19; 36), reference signals (12; 5), or both (33; 38; 14; 7; 30; 10; 3). In practice, these pipelines cover complementary objectives, including full reference-guided variation (14; 10), temporal control (5), localized modification such as inpainting (12) or region-specific editing (19). These developments are particularly relevant to SFX generation, where the objective is often not only to synthesize plausible audio, but also to preserve event identity and temporal structure while supporting controllable variation (5; 3). In addition, a few methods are explicitly designed to adapt to new domains under restricted-data settings, including retrieval-augmented approaches for zero-shot and few-shot TTA generation (34) and customized generation from a few reference audio samples (35). Yet these methods remain primarily generation-oriented and are not necessarily suited to editing or reference-guided SFX workflows. 2.3 Evaluation Gaps for Production-oriented SFX Variation SFX generation focuses on producing short, semantically meaningful, and often temporally structured events tied to an action, scene, or reference. (5; 3). Compared with general TTA generation, it places stronger emphasis on event identity, temporal synchronization, and controllability. Despite recent progress, such methods are rarely evaluated under a shared SFX-variation objective or a common benchmark, and their evaluation protocols remain tied to each method’s original task: AudioLDM mainly on TTA generation (11), with style transfer, high resolution, and inpainting shown separately; T-Foley on temporal-event control, combining quality, diversity, and temporal-adherence metrics with MOS ratings; ThinkSound on VTA generation, object-focused generation, and editing, mainly on VGGSound (2) and MovieGen Audio Bench (27); AudioX on a large protocol spanning TTA (31; 29), VTA, joint text-and-video, and instruction-driven benchmarks; and A2SB on localized restoration. Thus, evaluations reflect different task definitions, datasets, and control assumptions, making a direct cross-method comparison inherently difficult. This limitation is important for production-oriented SFX variation, where the objective is not only to generate plausible audio but also to preserve event identity, support controllable variation, and remain feasible under practical constraints such as limited data and efficient inference. Such limitations motivate a shared evaluation framework suitable for practical SFX deployment (this being the focus of our work). Figure 1: Overview of the proposed production-oriented evaluation framework for reference-guided SFX variation. A. In production settings, a reference sound often requires multiple diverse variations (Sec. 3.1); B. Existing baselines address complementary capabilities (17; 5; 19; 12; 28), including full variation generation (Sec. 3.2); C. We propose an evaluation framework to assess models’ abilities at generating variations of SFX in production settings (Sec. 3) and report a full capability profile for each baseline on our web page LABEL:fn:web. 3 Production-Oriented Evaluation Framework and Experimental Set-up In production settings, a reference sound often requires diverse variations under constraints such as identity preservation, controllability, and temporal alignment. Existing baselines address complementary capabilities, but are rarely compared under a shared objective. Figure 1 summarizes the proposed evaluation framework to standardize comparison through a common ESC-50 (25) reference-guided ATA setting complemented by method-specific analyses. Each method is then evaluated using objective, perceptual, and qualitative criteria to characterize its relevance to different SFX production needs. Table 1: Production requirements for reference-guided SFX variation, with evaluation signals and typical failure modes. Requirement Assessment Typical failure modes R1 Fidelity and realism FAD↓ , MOS↑ , spectrogram inspection, transient diagnosis Transient smearing, noise, phasiness, synthetic texture, unnatural reverberation R2 Identity preservation S-MOS↑ , ImageBind alignment↑ Semantic drift, generic textures, unrelated events R3 Diversity without drift ImageBind diversity↑ , pairwise variation analysis Near-duplicate outputs, limited variation, identity drift R4 Temporal alignment Energy-curve comparison, onset timing error, FWHM ratio Shifted onsets, stretched/compressed events, incorrect ordering R5 Energy control Energy-curve & envelope comparison, pre-onset Δ Loudness drift, distorted dynamics, poor envelope matching R6 Controllability Comparison across control settings Weak control response, uncontrolled drift, over-transformation R7 Targeted modification Local qualitative and editing analysis Global rewriting, boundary discontinuities, unintended off-target changes R8 Robustness and stability Cross-condition comparison, failure analysis Unstable behavior, out-of-domain failure, inconsistent quality R9 Efficiency Inference time and compute-cost comparison Excessive latency, slow sampling, impractical compute cost 3.1 Production Requirements Production-ready SFX variation requires more than fidelity: a useful method should preserve identity, support controlled variation, while remaining efficient for iterative deployment. In this work, the evaluation is structured around nine production requirements, each associated with an operational definition, specific evaluation signals, and typical failure modes; see Table 1. These requirements guide the baseline selection and the evaluation protocol. 3.2 Heterogeneous Baselines We evaluate five representative baselines covering complementary SFX variation capabilities. AudioLDM (17) and AudioX (28) for SFX morphing; T-Foley (5) for time-conditioned control; ThinkSound (19) for targeted editing; and A2SB (12) for inpainted variations. A detailed mapping between the production requirements and each baseline can be found on our web pageLABEL:fn:web. Below, we provide a brief description mapping of their conditioning signals, editing scope and production strengths. AudioLDM (17) can transfer the reference style to generate reference-conditioned variations. Its primary control mechanism is text conditioning, while its editing scope covers the entire generated sample. Thus, it serves as a general-purpose baseline for full-sample variation, quality and realism (R1), diversity (R3), and flexible user control through text or reference-guided editing (R6). T-Foley (5) uses a protocol centered on temporal-event control through sound-class and RMS envelope information for Foley generation. This makes T-Foley particularly relevant for temporal alignment (R4) and energy control (R5), while supporting identity preservation (R2). However, it does not support targeted local edits of a specific waveform region (R7), since the reference is used only as an event condition rather than as a direct ATA edition. ThinkSound’s (19) control mechanism combines semantic instructions with reference conditioning, enabling targeted modifications while preserving overall plausibility. ThinkSound supports a localized editing scope, making it suitable for studying controllable semantic variation and localized content modification. This allows evaluation of targeted modification (R7), controllability (R6), and semantic editing while maintaining plausible variation (R3). AudioX (28) combines text, video, and audio signals through multimodal fusion, enabling audio generation under multiple conditioning signals. AudioX spans TTA, VTA, joint text-and-video conditioning, and instruction-driven conditioning. This is particularly relevant as a recent generation model with strong identity preservation (R2), diversity without drift (R3), and controllable reference-guided generation (R6), while also providing a useful reference point for fidelity and realism (R1). A2SB (12) masks a segment and reconstructs it from the surrounding unmasked context. Unlike the other baselines, control is provided directly through the preserved waveform context rather than text or temporal conditioning. Its editing scope is strictly local, affecting only the masked region while leaving the remainder of the signal unchanged. This makes A2SB relevant for evaluating targeted modification (R7), the preservation of an unedited context as part of identity preservation (R2), and localized variation with limited drift (R3). Overall, the selected baselines cover four distinct editing scopes: (i) full-sample generation (AudioLDM, AudioX), (i) temporally constrained generation (T-Foley), (i) semantic reference-guided editing (ThinkSound), and (iv) localized waveform inpainting (A2SB). This diversity enables a broader assessment of production-oriented SFX variation beyond fidelity alone. 3.3 Evaluation Framework Task Formulation — We evaluate production-oriented SFX variation through a two-stage protocol after lightweight adaptation on the few-shot soundbank ESC-50 (25): (i) ATA generation: Given a reference audio clip and its corresponding class label, each model generates variants, which are then evaluated comprehensively. (i) Method-specific analysis: Selected methods are further analyzed in their primary audio editing setting to highlight their strengths, including SFX morphing, inpainting, temporal alignment, and targeted editing. Dataset — Experiments use ESC-50, with 2,000 five-second recordings across 50 classes. We use the official split, with folds 1-3 for training, fold 4 for validation, and fold 5 for testing, yielding 8 clips per class per fold. This low-data setting approximates practical adaptation of pretrained models to limited SFX soundbanks. Although relatively small in scale, the ESC-50 setting reflects practical deployment scenarios in which pretrained models are fine-tuned on limited subsets of target classes rather than trained from scratch on large-scale datasets. Fine-Tuning Protocol — Fine-tuning is performed on top of each model’s initial pretraining, with particular attention to preserving each model’s original setup and framework. All baselines are initialized from their released pretrained checkpoints. Fine-tuning is deliberately lightweight to adapt each method to ESC-50 while preserving its native behavior. Encoders are kept frozen when applicable, and only task-relevant generation or conditioning layers are adapted; method-specific details are reported in our web pageLABEL:fn:web. ATA Generation Protocol — For each of the 400 ESC-50 test references, we generate N = 10 variants, yielding 4,000 outputs per model. Each method receives the waveform reference and class name as minimal conditioning. ATA generation is implemented according to each baseline’s native design: T-Foley uses reference-derived temporal features and RMS-envelope conditioning; ThinkSound and AudioX sample from a noisy reference latent or waveform, respectively, with class-name conditioning; and A2SB is evaluated in its native inpainting regime by masking a segment of the reference and sampling multiple restorations from the same corrupted input. Table 2 reports inference time and number of generated variants per hour for each method, serving as a practical efficiency indicator (R9). Results in Tab. 2 are diagnostic rather than directly comparable, since methods use different backbones, batch sizes, and sampling steps. Method-specific analysis Protocol —Because the selected baselines differ in their conditioning signals and editing scope, we complement the shared ATA framework with a capability analysis on each method. These analyses are diagnostic and are not intended for direct cross-method comparison. For A2SB, we evaluate localized inpainting by comparing short masks of 0.3-1.0,s with longer masks of 0.3-2.0,s, using metrics computed exclusively on the inpainted regions. For T-Foley, we assess the effect of its restricted pretraining manifold (only seven classes, e.g., coughing, dog, footsteps, keyboard_typing, and rain) by separating ESC-50 classes that overlap with its native training domain from the remaining unseen classes. For AudioLDM and AudioX, we evaluate SFX morphing using reference-to-target transformations under different noise levels σ, which control the strength of the transformation. For ThinkSound, we evaluate object-centric editing on masked regions using attenuation, enhancement, and reverberation instructions. These native analyses preserve each model’s intended control mechanism while complementing the requirement-level interpretation of the shared framework. 3.4 Evaluation Metrics and Listening Study Objective evaluation is performed on cropped 4 s clips for all methods. Distributional audio quality and realism (R1) are measured with Fréchet Audio Distance (FAD) using the AudioLDM-eval implementation (16), computed per class and pooled across classes. Reference alignment and diversity — We use ImageBind audio embeddings (9) for identity preservation (R2) and diversity (R3), since the shared audio-to-audio (ATA) protocol is reference-conditioned rather than text-driven: an audio-text metric such as CLAP would mainly reflect agreement with the class label rather than preservation of the specific reference event. Let z^rref z^ref_r and z^r,vgen z^gen_r,v be the ℓ2 _2-normalised embeddings of reference r and its variant v. Alignment is the reference-variant cosine similarity Ar,v=cos_sim(z^ref,z^r,vgen)=⟨z^ref,z^r,vgen⟩‖z^ref‖2‖z^r,vgen‖2A_r,v=cos\_sim( z^ref_r, z^gen_r,v)= z^ref_r, z^gen_r,v \| z^ref_r\|_2\,\| z^gen_r,v\|_2 and diversity DrD_r is the mean pairwise cosine distance Dr=2Vr(Vr−1)∑1≤i<j≤Vr(dcos(z^r,igen,z^r,jgen))D_r= 2V_r(V_r-1) _1≤ i<j≤ V_r (d_ ( z^gen_r,i, z^gen_r,j ) ), where dcos(⋅,⋅)=1−cos_sim(⋅,⋅)d_ (·,·)=1-cos\_sim (·,· ) and Vr=10V_r=10 denotes the variants of a reference, following (32). High diversity must be read jointly with FAD and alignment, as a large spread may also indicate semantic drift. Transient diagnostics — Temporal behaviour (R4–R5) is assessed from a normalised onset-strength envelope e~t e_t, built by summing positive log-mel increments over mel bins and rescaling to [0,1][0,1], so the comparison reflects transient shape rather than loudness. Peaks are detected with a 40 ms minimum spacing and prominence ρ=0.10ρ=0.10, yielding three measures: (i) FWHM ratio, the generated-to-reference ratio of median peak full-width at half-maximum (→1→ 1 ideal; >1>1 smeared, <1<1 sharpened); (i) pre-onset Δ , the difference in normalised pre-transient energy (4040 ms before vs. 8080 ms after the peak; →0→ 0 ideal, positive values indicate pre-echo); and (i) onset error, the median timing gap between matched reference and generated onsets (tolerance δmax=150 _ =150 ms). All detection parameters are fixed across methods. Listening study — To assess perceptual identity (R1–R2), the Similarity Mean Opinion Score (S-MOS) is rated by 15 participants on a 1–5 identity scale over reference–variation pairs drawn from the 4,000 outputs per method. Participants used headphones in a quiet environment, could replay each pair freely, and rated 100 randomized trials per method with the prompt: rate the identity fidelity (1–5) as the similarity to the reference event/source, excluding loudness. Scores are reported with 95% CI. Table 2: Practical inference cost under each method’s setting for generating N=10N=10 variants per reference over 50 ESC-50 classes with 8 references per class (4 000 generations per model) (25). Method Backbone Hardware Steps Batch Inference Variants/ h ThinkSound (NeurIPS’25) MMDiT A100 40GB 24 2 0.66 h ∼ 6061 T-Foley (ICASSP’24) Wave. Diff. + FiLM RTX A6000 250 16 28.8 h ∼ 139 A2SB (arXiv’25) Attn. UNet A100 40GB 300 1 33 h ∼ 121 AudioX (ICLR’26) DiT A100 40GB 250 1 10 h 400 AudioLDM (ICML’23) UNet (LDM) RTX A6000 200 1 0.75 h ∼ 5333 4 Experimental Results Table 3: Overall quantitative and transient-level diagnosis comparison of baselines for ATA on SFX variation. FAD, S-MOS, diversity, and ImageBind-audio alignment summarize global generation quality, perceptual quality, variation diversity, and reference consistency. Transient diagnostics measure onset-local behavior using log-mel onset-strength envelopes: FWHM Ratio measures onset broadening or sharpening, Pre-onset Δ measures extra energy before the transient, and Onset Error evaluates temporal alignment. S-MOS is reported as the mean with 95% confidence intervals over 15 raters. Best and second best on full generation methods. Overall Quantitative Results Transient Level Diagnostics Methods FAD↓ S-MOS↑ Div.↑ Align.↑ FWHM Ratio →1→ 1 Pre-onset Δ →0→ 0 Onset Err. ms↓ AudioX (ICLR, 2026) 9.34 3.37 [3.08, 3.66] 0.27 0.59 0.985 -0.004 43.52 AudioLDM (ICML, 2023) 20.09 2.22 [1.97, 2.48] 0.43 0.39 0.969 -0.005 60.72 ThinkSound (NeurIPS, 2025) 16.51 2.57 [2.25, 2.89] 0.32 0.51 0.978 -0.011 53.01 T-Foley (ICASSP, 2024) 24.53 1.89 [1.62, 2.16] 0.32 0.22 1.011 0.008 63.53 A2SB∗ (ArXiv, NVIDIA, 2025) 4.43 4.81 [4.75, 4.87] 0.28 0.49 0.992 0.005 54.79 ∗ All reported objective metrics are computed on cropped 4 s clips for all methods, except for A2SB where only inpainted regions are evaluated. This section presents results of the proposed production-oriented evaluation framework for reference-guided SFX variation. We report the shared ATA evaluation, Pareto-inspired trade-offs, and qualitative analyses, followed by method-specific results on native capabilities not fully captured by the shared framework. 4.1 ATA Generation Table 3 reports the shared reference-conditioned ATA comparison along with transient-level diagnostics, further discussed in Sec. 4.3. We evaluate each method’s ability to generate variants from a reference SFX across reference preservation, quality, diversity, and perceptual identity under common production constraints. AudioX emerges as the strongest full-generation baseline, providing the most balanced trade-off across audio quality, identity preservation, and human evaluation. It remains consistently competitive across the main criteria, with a low FAD of 9.34, the strongest alignment score of 0.59, and the highest S-MOS among full-generation methods (3.37) [3.08, 3.66]. Its diversity score is more restrained than AudioLDM (0.27), suggesting that AudioX favors reference consistency over large variation, but without collapsing to over-fitting the reference. In contrast, AudioLDM reaches the highest diversity among full-generation baselines (0.43), but with weaker alignment (0.39), lower S-MOS (2.22 [1.97, 2.48]), and higher FAD (20.09), suggesting identity drift. ThinkSound presents an intermediate profile, with diversity (0.32), alignment (0.51), and S-MOS (2.57 [2.25, 2.89]), indicating that controlled variation does not fully translate into stronger perceptual fidelity in this ATA setting. Although T-Foley explicitly conditions on temporal structure and energy, it performs poorly in the shared ATA setting, with the highest full-generation FAD (24.53), the lowest S-MOS (1.89 [1.62, 2.16]), and weak alignment (0.22), despite diversity comparable to ThinkSound (0.32). This suggests that temporal-event conditioning alone is insufficient to preserve the broader perceptual identity of the reference. A2SB achieves the strongest S-MOS and FAD values, but should be interpreted separately because its variations are restricted to short inpainted regions of 0.3-1.0 s while most of the reference remains preserved. On inpainted regions, it obtains FAD (4.43), S-MOS (4.81 [4.75, 4.87]), diversity (0.28), and alignment (0.49), confirming its strength for localized repair rather than full ATA generation. Among the full-generation baselines evaluated under the shared ATA setting, results indicate that AudioX provides the strongest production-oriented compromise when identity preservation and perceptual fidelity are prioritized, while AudioLDM and ThinkSound expose different diversity-identity trade-offs that are further analyzed in Sec. 4.4. 4.2 Pareto-inspired Analysis in Production Settings Figure 2: Diversity–identity alignment (R3-R2) trade-off across reference-guided SFX variation methods. Each point represents a model positioned according to its ability to preserve reference identity and generate diverse variations. Higher diversity and alignment are better. Figure 2 summarizes the diversity-alignment trade-off AudioLDM provides the highest diversity (R3: 0.43), but weaker identity alignment (R2: 0.39), while AudioX reaches the strongest alignment (R2: 0.59) with lower diversity (R3: 0.27). ThinkSound lies between these two regimes, with R3: 0.32 and R2: 0.51, for a moderate balance between variation and reference consistency. T-Foley does not occupy a favorable Pareto position, since it has diversity comparable to ThinkSound (R3: 0.32), but much weaker alignment (R2: 0.22). The position of A2SB should again be interpreted separately, since its generation is restricted to a short inpainted segment rather than full-clip generation. Overall, the analysis confirms that higher diversity alone is insufficient when accompanied by weaker identity preservation 222Further Pareto-inspired analyses are provided in our web pageLABEL:fn:web. Figure 3: Qualitative comparison for the ATA variation task on the human-sound example laughing. For A2SB inpainting, the masked and regenerated sections between 0.3 s and 1 s are outlined. Similar texture to reference is outlined in white. Top: mel-spectrogram. Bottom: energy curve. 4.3 Further Qualitative Analysis Figure 3 provides a visual analysis of the spectro-temporal structure and frequency behavior of the baseline generations. While the quantitative results capture overall trends in quality, alignment, and diversity, the spectrogram and energy-curve visualizations reveal additional differences in local texture preservation, temporal organization, and failure modes across methods. In terms of spectral texture, AudioLDM shows the best-preserved fine local structure, as highlighted by the white boxes, whereas T-Foley generations exhibit noisier textures, attenuated high-frequency content, and a reduced spectral range. ThinkSound and AudioX remain closer to the reference at the event level, although ThinkSound shows more local spectral smoothing. The transient diagnostics support this observation: ThinkSound has an FWHM ratio close to 1 but a moderate onset error (53.01 ms), while AudioX achieves the lowest onset error among full-generation methods (43.52 ms), supporting stronger temporal organization. These differences also appear in the temporal-energy curves. T-Foley sometimes follows the coarse reference energy profile, which is expected from its explicit energy conditioning, but this does not translate into better onset-level alignment, since it has the largest onset error among the full-generation baselines (63.53 ms). This suggests that envelope following does not necessarily preserve sound texture. A2SB preserves the global energy evolution well because most of the reference remains unchanged, while local transient diagnostics on the inpainted region remain close to the reference, with a FWHM ratio of 0.992 and a pre-onset difference of 0.005. However, its inpainted regions may still introduce localized energy changes. Overall, the qualitative analysis complements the quantitative comparison by revealing method-specific trade-offs in local texture preservation, temporal organization, and energy behavior that are not fully captured by scalar metrics alone. 4.4 Method-specific Analysis Table 4 reports compact capability-specific diagnostics under each method’s native setting defined in Sec 3.3. These results complement the shared ATA benchmark by showing where each method is most suitable within its intended editing scope. Below, the key observations from each method are examined in greater detail. Table 4: Capability-specific diagnostics under each method’s native setting. Results are diagnostic and not intended for direct cross-method comparison. Methods FAD↓ S-MOS↑ Div.↑ Align.↑ A2SB Inpainting (evaluated on 3 classes) Mask 0.30.3–1.01.0 s 4.74 – 0.39 0.35 Mask 0.30.3–2.02.0 s 7.57 – 0.11 0.31 SFX morphing (evaluated on 3 classes) AudioX ↓σ σ / ↑σ σ 4.57 / 6.64 – 0.10 / 0.26 0.77 / 0.64 AudioLDM ↓σ σ / ↑σ σ 18.97 / 24.09 – 0.34 / 0.56 0.40 / 0.15 T-Foley Restricted pretraining manifold Seen classes (5 classes) 13.51 3.05± 0.78 0.34 0.34 Unseen classes (45 classes) 25.75 1.76± 0.62 0.32 0.21 ThinkSound Object-centric region editing (evaluated on 5 classes) Attenuation Mask 1.01.0–4.04.0 s 9.06 – 0.19 0.67 Enhancement Mask 1.01.0–4.04.0 s 9.30 – 0.21 0.64 Reverberation Mask 1.01.0–4.04.0 s 8.97 – 0.22 0.64 A2SB on inpainting — Table 4 shows that A2SB remains strong for local repair, but degrades as mask duration increases. Specifically, FAD rises from 4.744.74 to 7.577.57 and alignment stabilizes from 0.350.35 to 0.310.31, while diversity decreases from 0.390.39 to 0.110.11. This suggests that longer masked regions are harder to reconstruct and may reduce both consistency and variation, likely because less local context remains available to guide the reconstruction. The transient diagnostic in our web pageLABEL:fn:web further shows that both mask regimes remain temporally stable at the onset level, but a longer mask has a lower FWHM ratio (0.890.89), suggesting narrow or sharp reconstructed transients. T-Foley on temporal-energy control and restricted pretraining manifold — The noisy spectral texture visible in Figure 3, supports the quantitative results: although temporal-event conditioning can impose coarse structure, it does not guarantee identity preservation, and resulting variations show degraded fidelity and semantic consistency. This behavior is explained with T-Foley’s restricted pretraining manifold; Table 4 shows substantially better within-manifold performance, with FAD improving from 25.7525.75 to 13.5113.51, S-MOS from 1.76±0.621.76± 0.62 to 3.05±0.783.05± 0.78, and alignment from 0.210.21 to 0.340.34, while diversity remains stable (0.32→0.340.32→ 0.34). Together, these results suggest that T-Foley’s weaker ATA performance is driven more by fidelity, identity preservation, and domain-transfer limitations than by onset timing alone. AudioLDM and AudioX on SFX Morphing Task — Following the AudioLDM audio style-transfer setup (17) adapted to ESC-50, three examples are evaluated: toilet flush to children singing, sheep to narration/monologue, and coughing to ambient music. For smaller σ values, the outputs remain closer to the reference, whereas larger σ values increase adherence to the target text condition. Figure 4 shows the reference audio together with generated samples obtained at different initialization noise levels σ, which control the transfer strength. At low σ, AudioX preserves the source timbre and spectral structure more closely, consistent with its higher alignment (0.770.77 vs. 0.400.40) and lower diversity (0.100.10 vs. 0.340.34). As σ increases, both methods trade alignment for diversity, with sharper results for AudioLDM (alignment 0.40→0.150.40→ 0.15; while diversity increases 0.34→0.560.34→ 0.56). AudioX changes more progressively (alignment 0.77→0.640.77→ 0.64; diversity 0.10→0.260.10→ 0.26). The transient diagnostic in our web pageLABEL:fn:web follows the same tendency, with onset error increasing at higher σ for both methods. This yields a clear trade-off: AudioX provides smoother and more identity-preserving transformations, whereas AudioLDM produces stronger transformations at the cost of weaker preservation of the original audio identity. Figure 4: Ablations on SFX Morphing task for AudioLDM (17) and AudioX (28) on ESC-50 (25). From left to right: the reference audio (e.g., sheep, toilet_flush, cough) and four generated samples conditioned on the target text prompt with different initialization noise levels σ (transfer strength). ThinkSound on object-centric editing — A strength of this baseline is its ability to refine or modify sounds through user-specified regions and semantic instructions. Since the shared ATA protocol only provides restricted class-level conditioning, we further evaluate it in its native object-centric editing setting. Table 4 shows that the edited regions remain close to the reference, with alignment (0.640.64–0.670.67) and moderate diversity (0.190.19–0.220.22). Figure 5 visually supports these results through light but directionally consistent fluctuations. Attenuation slightly reduces the target transient regions, enhancement increases transient strength, and reverberation introduces more diffuse repeated patterns. However, the editing scope remains limited in this setting. Since the prompts are deliberately simple and the evaluated references often contain a single dominant sound event, the model has limited semantic or acoustic structure to selectively modify. Further, the transient diagnostic in our web pageLABEL:fn:web shows reasonable local timing preservation, with onset errors between 38.3 and 43.8 ms, although FWHM ratios above 1.1 suggest some transient broadening. As a result, ThinkSound produces plausible local changes, but the edits remain relatively subtle and do not always correspond to strong, fine-grained acoustic transformations. Figure 5: Ablation of ThinkSound on object-centric local editing over masked regions of the reference (e.g. crow class from ESC-50 (25)). Three instructions are evaluated: a. Attenuation, "reduce the class_name volume to sound far away and muffled, in the target region only."; b. Enhancement, "make class_name sound in the target region stronger and sharper."; and c. Reverberation, "make class_name event more distant and reverberant in the target region only.". Masked and edited sections [1 s; 3 s] are outlined. 5 Discussion and Conclusion Production-ready SFX variation requires an evaluation protocol able to compare heterogeneous generation and editing methods under shared production requirements. By combining a common reference-guided ATA task with capability-specific analyses, our framework makes these trade-offs explicit rather than reducing them to a single aggregate score. Under the shared reference-guided ATA generation setting, the baselines reveal complementary strengths across production needs. Among the full-generation methods, AudioX provides the strongest overall compromise between fidelity, identity preservation, diversity, and reference alignment. In contrast, AudioLDM remains more suitable for stronger stylization and higher variation, T-Foley for explicit temporal and energy control, A2SB for localized inpainting and repair, and ThinkSound for targeted editing when richer semantic guidance is available. Thus, from a production perspective, no single baseline dominates: AudioX offers the strongest overall compromise among full-generation methods, while the most suitable choice ultimately depends on workflow priorities. Heterogeneous baselines indicate that production-ready SFX variation can’t be assessed through generic audio-quality benchmarks alone. Our framework addresses this gap through requirement-driven baseline selection, a shared ATA framework, and complementary method-specific analyses, enabling a more structured comparison while preserving native baseline strengths. Several limitations remain, including heterogeneous pretrained setups and the restricted scope of ESC-50 relative to real production soundbanks. A suitable pipeline must jointly balance realism, identity preservation, controllable variation, and practicality. More broadly, this work points to future steps for model design: unified full-reference variation with explicit control, localized editing, and efficient deployment within one framework. Acknowledgments: This work was supported by the Natural Sciences and Engineering Research Council of Canada, with additional computational resources provided by the Digital Research Alliance of Canada. We thank Dr. Hugo Seuté for his valuable support, the La Forge R&D department and the Alice sound team at Ubisoft for contributing to the problem formulation and insights in defining the production requirements. References [1] M. A. I. Benjamin Elizalde and H. Wang (2023) CLAP learning audio concepts from natural language supervision. In ICASSP, External Links: Document Cited by: §6.2.1. [2] H. Chen, W. Xie, A. Vedaldi, and A. Zisserman (2020) VGGSound: a large-scale audio-visual dataset. In ICASSP, External Links: Document Cited by: §2.3. [3] Z. Chen, P. Seetharaman, B. Russell, O. Nieto, D. Bourgin, A. Owens, and J. Salamon (2025) Video-guided foley sound generation with multimodal controls. In CVPR, External Links: Document Cited by: §1, §2.1, §2.2, §2.3. [4] H. K. Cheng, M. Ishii, A. Hayakawa, T. Shibuya, A. Schwing, and Y. Mitsufuji (2025) MMAudio: taming multimodal joint training for high-quality video-to-audio synthesis. In CVPR, Cited by: §2.1, §6.2.1, §6.2.1. [5] Y. Chung, J. Lee, and J. Nam (2024) T-FOLEY: a controllable waveform-domain diffusion model for temporal-event-guided foley sound synthesis. In ICASSP, External Links: Document Cited by: §1, §1, Figure 1, §2.1, §2.2, §2.3, §3.2, §3.2, Table 7, Table 7, Table 7, §6.3. [6] Z. Evans, J. D. Parker, C. Carr, Z. Zukowski, J. Taylor, and J. Pons (2025) Stable audio open. ICASSP. External Links: Document Cited by: §1, §1. [7] P. Fang, Y. He, Y. Xing, Q. Chen, S. Lim, and H. Yang (2026) AC-foley: reference-audio-guided video-to-audio synthesis with acoustic transfer. In ICLR, Cited by: §1, §2.2, Table 7, Table 7. [8] H. F. García, O. Nieto, J. Salamon, B. Pardo, and P. Seetharaman (2025) Sketch2Sound: controllable audio generation via time-varying signals and sonic imitations. In ICASSP, Cited by: §2.1. [9] R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra (2023) ImageBind: one embedding space to bind them all. In CVPR, Cited by: §3.4, Table 7, Table 7, §6.2.1, §6.2.1. [10] Y. Jia, Y. Chen, J. Zhao, S. Zhao, W. Zeng, Y. Chen, and Y. Qin (2024) AudioEditor: a training-free diffusion-based audio editing framework. arXiv preprint arXiv:2409.12466. Cited by: §2.2. [11] C. D. Kim, B. Kim, H. Lee, and G. Kim (2019) AudioCaps: generating captions for audios in the wild. In NAACL-HLT, p. 119–132. Cited by: §2.3. [12] Z. Kong, K. J. Shih, W. Nie, A. Vahdat, S. Lee, J. F. Santos, A. Jukic, R. Valle, and B. Catanzaro (2025) A2SB: audio-to-audio schrodinger bridges. arXiv. Cited by: §1, Figure 1, §2.2, §3.2, §3.2, Table 7, Table 7, Table 7, Table 7, §6.3. [13] F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. Défossez, J. Copet, D. Parikh, Y. Taigman, and Y. Adi (2023) AudioGen: textually guided audio generation. In ICLR, Cited by: §1. [14] Z. Lan, Y. Hao, and M. Zhao (2026) SmartDJ: declarative audio editing with audio language model. In ICLR, Cited by: §1, §2.2, Table 7. [15] J. Liang, Y. Chen, Y. Yuan, D. Jia, X. Zhuang, Z. Chen, Y. Wang, and Y. Wang (2025) AudioMorphix: training-free audio editing with diffusion probabilistic models. arXiv. Cited by: §1, Table 7, Table 7, Table 7. [16] H. Liu, Y. Zhang, R. Mira, Y. Zang, and I. E. Ashimine (2023)audioldm_eval: audio generation evaluation(Website) Note: GitHub repository Cited by: §3.4. [17] H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley (2023) AudioLDM: text-to-audio generation with latent diffusion models. In ICML, Cited by: §1, Figure 1, §2.1, §3.2, §3.2, Figure 4, §4.4, Table 7, Table 7, §6.3. [18] H. Liu, Y. Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y. Wang, W. Wang, Y. Wang, and M. D. Plumbley (2024) AudioLDM 2: learning holistic audio generation with self-supervised pretraining. ACM TASLP. External Links: Document Cited by: §1. [19] H. Liu, K. Luo, J. Wang, W. Wang, Q. Chen, Z. Zhao, and W. Xue (2025) ThinkSound: chain-of-thought reasoning in multimodal large language models for audio generation and editing. In NeurIPS, Cited by: §1, §1, Figure 1, §2.1, §2.2, §3.2, §3.2, Table 7, Table 7, Table 7, Table 7, Table 7, §6.2.1, §6.2.1, §6.3. [20] Y. Liu, C. Jin, and D. Gunawan (2023) DDSP-sfx: acoustically-guided sound effects generation with differentiable digital signal processing. DAFx. Cited by: §2.1. [21] S. Luo, C. Yan, C. Hu, and H. Zhao (2023) Diff-foley: synchronized video-to-audio synthesis with latent diffusion models. In NeurIPS, Cited by: §1. [22] N. Majumder, C. Hung, D. Ghosal, W. Hsu, R. Mihalcea, and S. Poria (2024) Tango 2: aligning diffusion-based text-to-audio generations through direct preference optimization. In ACM ICM, External Links: Document Cited by: §2.1. [23] X. Niu, J. Zhang, and C. P. Martin (2024) SoundMorpher: perceptually-uniform sound morphing with diffusion model. Cited by: §1. [24] Y. Okamoto, K. Imoto, S. Takamichi, R. Yamanishi, T. Fukumori, and Y. Yamashita (2021) Onoma-to-wave: environmental sound synthesis from onomatopoeic words. APSIPA Transactions. Cited by: §2.1. [25] K. J. Piczak (2015) ESC: dataset for environmental sound classification. In ACM M, External Links: Document Cited by: §1, §3.3, Table 2, §3, Figure 4, Figure 5. [26] B. Shi, A. Tjandra, J. Hoffman, H. Wang, Y. Wu, L. Gao, J. Richter, M. Le, A. Vyas, S. Chen, C. Feichtenhofer, P. Dollár, W. Hsu, and A. Lee (2025) SAM audio: segment anything in audio. arXiv. External Links: Document Cited by: §2.1. [27] T. M. G. team @Meta (2024) Movie gen: a cast of media foundation models. CVPR. Cited by: §2.3. [28] Z. Tian, Z. Liu, Y. Jin, R. Yuan, L. Xue, X. Tan, Q. Chen, W. Xue, and Y. Guo (2026) AudioX: diffusion transformer for anything-to-audio generation. In ICLR, Cited by: §1, Figure 1, §2.1, §3.2, §3.2, Figure 4, Table 7, Table 7, §6.2.1, §6.2.1, §6.3, §6.3. [29] H. Wang, C. Liu, J. Chen, H. Liu, Y. Jia, S. Zhao, J. Zhou, H. Sun, H. Bu, and Y. Qin (2026) TTA-bench: a comprehensive benchmark for evaluating text-to-audio models. AAAI. Cited by: §2.3. [30] Y. Wang, Z. Ju, X. Tan, L. He, Z. Wu, J. Bian, and S. Zhao (2023) AUDIT: audio editing by following instructions with latent diffusion models. In NeurIPS, Cited by: §2.2. [31] Z. Xie, X. Xu, Z. Wu, and M. Wu (2025) AudioTime: a temporally-aligned audio-text benchmark dataset. ICASSP. Cited by: §2.3. [32] Y. Xing, Y. He, Z. Tian, X. Wang, and Q. Chen (2024) Seeing and hearing: open-domain visual-audio generation with diffusion latent aligners. In CVPR, External Links: Document Cited by: §3.4, §6.2.1, §6.2.1. [33] D. Yang, J. Tian, X. Tan, R. Huang, S. Liu, X. Chang, J. Shi, S. Zhao, J. Bian, Z. Zhao, X. Wu, and H. Meng (2024) UniAudio: an audio foundation model toward universal audio generation. Cited by: §2.2. [34] M. Yang, B. Shi, M. Le, W. Hsu, and A. Tjandra (2025) AudioBox tta-rag: retrieval-augmented generation for zero-shot and few-shot text-to-audio. Arxiv. Cited by: §2.2. [35] Y. Yuan, X. Liu, H. Liu, X. Kang, Z. Chen, Y. Wang, M. D. Plumbley, and W. Wang (2025) DreamAudio: few-shot customized text-to-audio generation via natural language supervision. arXiv. Cited by: §2.2. [36] Y. Zhang, A. Maezawa, G. Xia, K. Yamamoto, and S. Dixon (2023) Loop copilot: conducting ai ensembles for music generation and iterative editing. Cited by: §2.2. [37] L. Zhao et al. (2025) UniForm: a unified multi-task diffusion transformer for audio-video generation. arXiv. External Links: Document Cited by: §2.1. [38] L. Zhao, L. Feng, D. Ge, R. Chen, F. Yi, C. Zhang, X. Zhang, and X. Li (2024) Prompt-guided precise audio editing with diffusion models. In PMLR, Cited by: §2.2. [39] G. Zhu, Y. Wen, M. Carbonneau, and Z. Duan (2023) EDMSound: spectrogram based diffusion models for efficient and high-quality audio synthesis. NeurIPS workshop Machine Learning for audio. Cited by: §1. 6 Appendix All material described in this appendix is available on the accompanying web pageLABEL:fn:web. 6.1 Capability profile of evaluated baselines Table 5: Capability profile on production requirements for reference-guided SFX variation across representative audio generation and editing methods. Color encodes each method’s suitability per requirement: native strength, supported / limited, weak / not supported. Further details on webpage. AudioX AudioLDM ThinkSnd. T-Foley A2SB R1 Fidelity & realism !50 !50 !50 !50 !50 R2 Identity preservation !50 !50 !50 !50 !50 R3 Diversity without drift !50 !50 !50 !50 !50 R4 Temporal alignment !50 !50 !50 !50 !50 R5 Energy control !50 !50 !50 !50 !50 R6 Controllability !50 !50 !50 !50 !50 R7 Targeted modification !50 !50 !50 !50 !50 R8 Robustness & stability !50 !50 !50 !50 !50 R9 Efficiency !50 !50 !50 !50 !50 6.2 Additional details on Evaluated Framework Table 6: Capability profile on production requirements for reference-guided SFX variation across representative audio generation and editing methods. Scores summarize the main evidence from the shared ATA benchmark and capability-specific analyses; local editing and inpainting results are diagnostic and not directly comparable to full-generation settings. Requirement AudioX (ICLR, 2026) AudioLDM (ICML, 2023) ThinkSound (NeurIPS, 2025) TFoley (ICASSP, 2024) A2SB (ArXiv, NVIDIA, 2025) R1 Fidelity & realism Strong full generation; low FAD (9.34) and best full-generation S-MOS (3.37) Limited; high FAD (20.09) despite plausible local texture Moderate; FAD (16.51), stronger in local editing FAD (8.97–9.30) Weak; highest full-generation FAD (24.53) and noisy spectra Strong local; FAD (4.43) on inpainted regions, not full generation R2 Identity preservation Strong; best full-generation alignment (0.59) and S-MOS (3.37) Limited; lower alignment (0.39) and S-MOS (2.22) Moderate; alignment (0.51) in ATA, higher local alignment (0.64–0.67) Weak; lowest alignment (0.22) and S-MOS (1.89) Strong local; S-MOS (4.81) and preserved context, but localized setting R3 Diversity without drift Balanced; diversity (0.27) with strongest alignment (0.59) High diversity / drift; best diversity (0.43) but weaker alignment (0.39) Moderate; diversity (0.32) with alignment (0.51); local edits remain subtle Drift-prone; diversity (0.32) but weak alignment (0.22) Localized; diversity (0.28) in ATA, mask-size sensitive in native analysis R4 Temporal alignment Strongest full generation; lowest onset error (43.52 ms), FWHM (0.985) Limited; onset error (60.72 ms), FWHM (0.969) Moderate; onset error (53.01 ms); native edits (38.3–43.8 ms) Mixed; explicit temporal conditioning but onset error (63.53 ms) Moderate local; onset error (54.79 ms), FWHM (0.992) R5 Energy control Implicit; preserves temporal organization but no explicit energy control Limited; no explicit energy condition Prompt-based; attenuation/enhancement/reverb produce light local fluctuations Native strength; RMS-envelope conditioning for temporal-energy control Context-preserving; unmasked energy mostly preserved, not explicit control R6 Controllability Strong; smooth SFX morphing with σ control, Align. (0.77 → 0.64) Strong but less stable; σ increases variation, Div. (0.34 → 0.56), Align. (0.40 → 0.15) Local but limited; semantic edits are directionally consistent but subtle under simple instructions Limited; class and RMS controls, no local semantic editing Native mask control; editable region/duration, but longer masks degrade FAD (4.74 → 7.57) R7 Targeted modification Limited; full variation / morphing rather than local target editing Limited; global style or morphing, no localized edit Native strength; object-centric edits with FAD (8.97–9.30), but limited fine-grained acoustic control Not supported; temporal generation, no waveform-region edit Native strength; masked inpainting on 0.3–1.0 s regions R8 Robustness & stability Good; stable morphing trade-off, FAD (4.57 → 6.64) Sensitive; stronger σ causes identity drop, Align. (0.40 → 0.15) Stable but subtle; local edits keep Align. (0.64–0.67), but effect size depends on instruction and sound complexity Domain-sensitive; seen vs. unseen FAD (13.51 → 25.75) Mask-sensitive; longer masks reduce performance, Div. (0.39 → 0.11) R9 Efficiency Moderate; 10 h for 4k generations, ∼ 400 variants/h Fast; 0.75 h, ∼ 5333 variants/h Fastest; 0.66 h, ∼ 6061 variants/h Slow; 28.8 h, ∼ 139 variants/h Slowest; 33 h, ∼ 121 variants/h Table 7: Production requirements for reference-guided SFX variation. Each requirement is paired with its operational definition, evaluation signals, reported evidence, and typical limitations or failure modes. ID Requirement Operational definition Evaluation signals / diagnostics Reported in Limitations and typical failure modes R1 Fidelity & realism Preserve transient and spectral details while remaining perceptually plausible and free of obvious artifacts as a real-world SFX. FAD↓ (plausibility-quality), MOS↑ , S-MOS↑ , spectrogram inspection (17; 28) Table 3, Fig. 6.4 and Fig. 3 FAD is distributional and should not be interpreted as direct perceptual fidelity. Typical failures include transient smearing, noise, phasiness, synthetic texture, and unnatural reverberation. R2 Identity preservation Preserve the reference event class and perceptual identity (7; 15). S-MOS↑ , ImageBind audio-reference alignment↑ (9) Table 3 Embedding similarity captures semantic alignment but not fine acoustic details. Typical failures include semantic drift, generic textures, and unrelated events. R3 Diversity without drift Generate plausible variations in texture, environment, or rendering while preserving event identity (7; 12; 15). ImageBind diversity↑ (9), pairwise ImageBind distance, pairwise variation analysis Fig. 2 High diversity may also reflect semantic drift. Typical failures include near-duplicate outputs, limited variation, and identity drift. R4 Temporal alignment Follow target temporal structure or explicit timing cues, including onset pattern, event order, and envelope shape (5; 19). Energy-curve comparison (5), envelope correlation, onset error, DTW on energy curves, temporal analysis (19) Fig. 6.4, Fig. 3 and Fig. LABEL:fig:Audio_energycurve_a2a, additional diagnostics in Sec. 4.3 and in Tab. 8 Temporal alignment is not directly measured by ImageBind similarity. Typical failures include shifted onsets, stretched or compressed events, and incorrect ordering. R5 Energy control Match, preserve, or controllably modify a target loudness profile or energy trajectory (5). Energy-curve analysis, RMS envelope comparison, energy-curve correlation, envelope comparison Fig. 6.4, Fig. 3 and Fig. LABEL:fig:Audio_energycurve_a2a, additional diagnostics in Sec. 4.3 and in Tab. 8 Energy curves provide only a coarse signal-level proxy. Typical failures include loudness drift, distorted dynamics, and poor envelope matching. R6 Controllability Modulate the strength or nature of the transformation through explicit user or model controls (28; 19; 14). Control-based comparison across settings, noise-level σ sweep, edit-strength analysis Fig. 4, Fig. 5 Control variables are method-specific and not always directly comparable. Typical failures include weak control response, uncontrolled drift, and over-transformation. R7 Targeted modification Modify specific segments or attributes while preserving the rest of the signal (15; 19; 12). Local qualitative analysis, editing evaluation, masked-region metrics, boundary analysis, local spectral distance A2SB / ThinkSound diagnostics Localized editing is not directly comparable to full-clip generation. Typical failures include global rewriting, boundary discontinuities, and unintended off-target changes. R8 Robustness & stability Maintain consistent behavior across challenging references, controls, mask sizes, domain shifts, or generation conditions (12; 19). Cross-condition comparison, failure analysis, mask-duration analysis, in-/out-manifold comparison Table 4 Diagnostic only; does not cover all real production conditions. Typical failures include unstable behavior, out-of-domain failure, and inconsistent quality. R9 Efficiency Remain practical for iterative production use in runtime, sampling cost, and deployment constraints (17; 12). Inference-time comparison, runtime, sampling steps, batch size, hardware, sample rate, cost discussion Table 2 Values are deployment-oriented and not fully normalized across implementations. Typical failures include excessive latency, slow sampling, and impractical compute cost. 6.2.1 Objective Evaluation Following prior audio-generation evaluations, ImageBind (9) is used for alignment (28; 19; 4) and diversity. This choice is motivated by the shared ATA generation protocol, which is centered on reference-conditioned audio variation rather than text-prompt adherence. Since all methods receive only restricted class-level text, an audio-text metric such as CLAP (1) may mainly reflect agreement with the class label rather than preservation of the specific reference event. In contrast, ImageBind allows direct audio-to-audio comparison in a common multimodal embedding space. ImageBind Alignment and Diversity — We compute audio embeddings using ImageBind (9) audio encoder; zrrefz^ref_r and zr,vgenz^gen_r,v are the embeddings of reference r and its generated variants v. First, embeddings are l2l_2-normalised. Alignment Ar,vA_r,v is computed as the cosine similarity (28; 19; 4) between each variant v and its reference r: Ar,v=cos_sim(z^ref,z^r,vgen)=⟨z^ref,z^r,vgen⟩‖z^ref‖2‖z^r,vgen‖2.A_r,v=cos\_sim( z^ref_r, z^gen_r,v)= z^ref_r, z^gen_r,v \| z^ref_r\|_2\,\| z^gen_r,v\|_2. (1) Although ImageBind is mainly used for alignment, we follow the Seeing and Hearing (32) method to estimate diversity through the ImageBind-space distance. Diversity DrD_r is computed for each reference as the average pairwise semantic distance between normalised embeddings of its generated variants: Dr=2Vr(Vr−1)∑1≤i<j≤Vr(dcos(z^r,igen,z^r,jgen))=2Vr(Vr−1)∑1≤i<j≤Vr(1−cos_sim(z^r,igen,z^r,jgen))=2Vr(Vr−1)∑1≤i<j≤Vr(1−⟨z^r,igen,z^r,jgen⟩‖z^r,igen‖2‖z^r,jgen‖2). splitD_r&= 2V_r(V_r-1) _1≤ i<j≤ V_r (d_ ( z^gen_r,i, z^gen_r,j ) )\\ &= 2V_r(V_r-1) _1≤ i<j≤ V_r (1-cos\_sim ( z^gen_r,i, z^gen_r,j ) )\\ &= 2V_r(V_r-1) _1≤ i<j≤ V_r (1- z^gen_r,i, z^gen_r,j \| z^gen_r,i\|_2\,\| z^gen_r,j\|_2 ). split (2) with the number of variants per reference Vr=10V_r=10 and ‖z^r,igen‖2‖z^r,jgen‖2=1\| z^gen_r,i\|_2\,\| z^gen_r,j\|_2=1 since embeddings are already l2l_2-normalized. To compute diversity, we follow prior work (32), which measures semantic distance in ImageBind space as d(x,y)=1−cos_sim(x,y)d(x,y)=1-cos\_sim(x,y). We extend the same distance to measure pairwise cosine distance among generated variants. This allows quantifying their spread among variations in the ImageBind embedding space. This metric should be interpreted jointly with FAD and reference-alignment scores. While diversity captures the spread among generated variants, high diversity alone may also reflect semantic drift. Combining these metrics provides a more balanced view of generation quality, identity preservation, and diversity among generated variations. To sum up, ImageBind alignment measures reference-variant similarity, while ImageBind diversity measures variation among generations from the same reference. These metrics capture semantic alignment and embedding-level diversity, whereas temporal synchronization and transient preservation are evaluated separately using signal-level onset diagnostics. Transient-level Diagnosis is computed through onset-level descriptors from the log-mel spectrogram of each reference and generated clip. Let ℓm,t _m,t denote the log-mel spectrogram, where m is the mel-frequency bin and t is the time frame. We first build a normalized onset-strength envelope in order to focus on transient shape rather than absolute loudness. We keep only positive temporal increases in the log-mel spectrogram: Dm,t=max(ℓm,t+1−ℓm,t,0),D_m,t= ( _m,t+1- _m,t,0), (3) and obtain the raw onset-strength envelope by summing over mel bins: et=∑mDm,t.e_t= _mD_m,t. (4) The envelope is then normalized to [0,1][0,1]: e~t=et−mintetmaxt(et−mintet)+ϵ. e_t= e_t- _te_t _t(e_t- _te_t)+ε. (5) This normalization makes the comparison more sensitive to onset shape and timing than to global energy differences. Transient peaks p are detected from e~t e_t using a minimum peak distance of 4040 ms and a peak prominence threshold ρ=0.10ρ=0.10. From these peaks, we report three diagnostic metrics. Full Width at Half Maximum (FWHM) ratio — For each detected peak p, the full width at half maximum is defined as: FWHM(p)=ωpΔms,FWHM(p)= _p _ms, (6) where ωp _p is the peak width in frames and Δms _ms is the frame duration in milliseconds. For each clip, we summarize FWHM(p)FWHM(p) by taking the median over all detected peaks. The generated-reference FWHM ratio is then: RFWHM=FWHMgenFWHMref+ϵ.R_FWHM= FWHM_genFWHM_ref+ε. (7) A value close to 11 indicates better transient-width preservation. Values above 11 suggest broader or more smeared transients, while values below 11 indicate sharper or narrower transients than the reference. Pre-onset Δ — For each peak p, we compare the normalized onset energy before and after the transient: Rpre(p)=∑t=max(0,p−qpre)p−1e~t∑t=pmin(T,p+qpost)−1e~t+ϵ,R_pre(p)= _t= (0,p-q_pre)^p-1 e_t _t=p (T,p+q_post)-1 e_t+ε, (8) where qpreq_pre and qpostq_post correspond to 4040 ms and 8080 ms windows, respectively. For each clip, Rpre(p)R_pre(p) is summarized by the median over detected peaks. The reported pre-onset difference is: Δpre=Rpre,gen−Rpre,ref. _pre=R_pre,gen-R_pre,ref. (9) Values close to 0 indicate similar pre-onset behavior. Positive values indicate extra energy before the transient, which may correspond to pre-echo or early energy leakage. Onset error — Let Tref=τirefT_ref=\ _i^ref\ and Tgen=τjgenT_gen=\ _j^gen\ be the detected onset times of the reference and generated clip. For each reference onset, we match the nearest unused generated onset: j∗(i)=argminj∈U|τjgen−τiref|,j^*(i)= _j∈ U | _j^gen- _i^ref |, (10) where U is the set of unmatched generated onsets. A match is accepted only if the distance is below the tolerance threshold δmax=150 _ =150 ms. For accepted matches, the onset error is; Eonset=1000⋅mediani|τj∗(i)gen−τiref|.E_onset=1000·median_i | _j^*(i)^gen- _i^ref |. (11) Lower values indicate better temporal alignment between reference and generated events. All transient-detection parameters are fixed across methods to ensure a consistent diagnostic protocol. We set the minimum peak distance to 40ms40\,ms to avoid counting multiple fluctuations from the same transient as separate onsets, and use a peak prominence threshold of ρ=0.10ρ=0.10 to suppress weak noise peaks in the normalized onset envelope. For the pre-onset diagnostic, we compare a 40ms40\,ms window before the peak with an 80ms80\,ms post-onset window, capturing short energy leakage relative to the local event energy. For onset matching, we use a tolerance of δmax=150ms _ =150\,ms, which allows moderate timing deviations while avoiding matches between unrelated events. These values are fixed analysis parameters and are not tuned per method. 6.2.2 Subjective Evaluation Further details on the subjective evaluation are provided as follows. All participants were instructed to perform the listening tests with headphones in a quiet environment. Each reference and variation pair could be replayed as many times as needed by the participant. Trials are anonymized and randomized, and scores are reported with 95% confidence intervals. For each method, the evaluation trials consisted of 100 test sets randomly sampled from a total of 4000 reference-variation test pairs. Each test set used the same evaluation prompt: "For each trial: listen to the Reference, then the Candidate. Rate the identity fidelity (1–-5), defined as the similarity to the reference event/source (excluding loudness)." 6.3 Additional Training Details All baselines are initialized from their original pretrained checkpoints: AudioLDM from audioldm-m-full (17), T-Foley from the released pretrained model (5), ThinkSound from the original checkpoints (19), AudioX from the HuggingFace release (28), and A2SB from the released masking-split checkpoints (12). For latent diffusion pipelines, e.g., AudioLDM and ThinkSound, the text and audio encoders are frozen, and only the diffusion backbone and conditioning projection layers are adapted. For T-Foley, semantic conditioning is kept fixed, and only the class embeddings, MLP embeddings, and FiLM conditioning layers are updated. For AudioX, fine-tuning follows the original stable-audio-tools pipeline, with ESC-50 clips zero-padded to the model’s fixed 11 s window and conditioning restricted to the class name and the noised reference clip. For A2SB, fine-tuning follows the original inpainting setup. The ESC-50 test fold contains 50 classes, with 8 reference clips per class. For each reference, N=10N=10 variants are generated, yielding 4 000 outputs per model. In the shared ATA generation setting, each model receives the waveform reference clip together with the class name as minimal textual conditioning. Latent diffusion pipelines — For AudioLDM and ThinkSound, the text and audio encoders are frozen to avoid overly sharp adaptation in the few-shot setting. Only the diffusion backbone (UNet or MMDiT) and the conditioning projection layers that map conditioning embeddings into the backbone channels are fine-tuned. For ThinkSound, conditioning features are pre-extracted into a latent-directory dataset, whereas for AudioLDM the conditioning layers are explicitly adapted. ESC-50 is converted into a latent-directory dataset before feature extraction and training. ThinkSound is fine-tuned for 10 epochs of 150 steps, while AudioLDM is fine-tuned for 200 training steps. The number of updates is intentionally kept low to reduce overfitting and distribution drift. Waveform diffusion pipeline — For T-Foley, fine-tuning is reduced from the original 500 epochs to 25 epochs with 250 steps per epoch, for a total of 6,250 training steps. T-Foley operates at 22 kHz. Since fine-tuning starts from pretrained checkpoints, the original sampling rate is preserved to avoid distribution mismatch. AudioX adaptation — AudioX is fine-tuned for 20 epochs with 200 steps per epoch, for a total of 4,000 training steps, using the original stable-audio-tools training pipeline (28). Because AudioX operates on a fixed 11 s window, each 5 s ESC-50 clip is zero-padded to 11 s and trained with a padding-mask loss, so that the diffusion objective is computed only on the real unpadded region. During training, conditioning uses only the class name of the clip, while the optional audio and video prompt modalities are provided as empty inputs to satisfy the model interface. A2SB adaptation and masking setup — For A2SB, two fine-tuning runs are performed to match the pretrained masking splits and support evaluation in the 0.3-1.0 s masking regime. One run is initialized from the 0.0-0.5 s checkpoint and the other from the 0.5-1.0 s checkpoint. To further probe robustness and extend editing ability, additional fine-tuning is performed under increasing masked durations. Two masking regimes are considered: short masked segments between 0.3 and 1.0 s, and longer masked segments between 0.3 and 2.0 s within the reference audio. For fine-tuning on ESC-50, waveforms are converted into STFT features following the original pipeline setup, with n_fft=2048, hop_length=512, and a sampling rate of 44.1 kHz. An inpainting corruption is then applied by masking a random noisy time segment, and the model is trained to reconstruct a clean STFT from the corrupted reference. Sampling-rate considerations — ThinkSound operates at 44 kHz, T-Foley at 22 kHz, and A2SB at 44.1 kHz in its original representation. Since all methods are adapted from pretrained checkpoints, their native sampling configurations are preserved during fine-tuning rather than forcing a common 16 kHz training setup, which could introduce distribution mismatch. Resampling to 16 kHz is applied only at evaluation time when required for fair comparison or listening tests. Figure 6: Pareto-Inspired Analysis across reference-guided SFX variation methods. We claim for pareto-inspired as for b) methods may not be direct equals in inference efficiency due to different backbones, batch size and audio representations. For all evaluated values; higher is better. 6.4 Additional Results and Evaluation Table 7 further details the production requirements, their associated evaluation signals and diagnostics, and where they are addressed in the paper. To complement this mapping, Table 6 provides a compact summary of each method’s capability profile, highlighting how well each method fulfills the requirements for SFX variation generation. Figures 6.4 and LABEL:fig:Audio_energycurve_a2a provide additional qualitative evidence for the shared ATA variation task. They complement the main quantitative comparison in Figure 3, by visualizing how each baseline preserves or alters spectro-temporal texture and energy behavior across further representative classes, such as crow, birds_chirping, handsaw. In Figure 6, the Pareto-inspired analysis shows that different methods occupy distinct production trade-offs. For identity preservation vs. temporal alignment, AudioX and A2SB define the frontier: AudioX provides the strongest temporal faithfulness, while A2SB achieves the highest identity preservation. However, this observation should be taken with caution, as the listening test S-MOS considers the full clip, while A2SB is an inpainting method and preserve unchanged regions of the reference. For identity preservation vs. efficiency, the frontier includes ThinkSound and AudioX, moving from high-throughput generation to higher identity preservation. Table 8: Capability-specific transient diagnostics under each method’s native setting. Results are diagnostic and not intended for direct cross-method comparison. Methods FWHM Ratio →1→ 1 Pre-onset Δ ↓ Onset Err. ms↓ A2SB Inpainting (evaluated on 3 classes) Mask 0.30.3–1.01.0 s 0.97 -0.02 55.07 Mask 0.30.3–2.02.0 s 0.89 0.01 51.99 SFX morphing (evaluated on 3 classes) AudioX ↓σ σ / ↑σ σ 0.95 / 0.87 0.007 / -0.02 58.61 / 67.78 AudioLDM ↓σ σ / ↑σ σ 0.95 / 1.03 -0.02 / -0.03 64.82 / 86.78 T-Foley Restricted pretraining manifold In-manifold (5 classes) 0.97 0.02 65.28 Out-manifold (45 classes) 1.01 0.007 63.33 ThinkSound Object-centric region editing (evaluated on 5 classes) Attenuation Mask 1.01.0–4.04.0 s 1.13 0.013 39 Enhancement Mask 1.01.0–4.04.0 s 1.15 0.01 38.3 ReverberationMask 1.01.0–4.04.0 s 1.12 0.01 43.8 black Animals (crow) Natural Sounds (chirping birds) Exterior Sounds (handsaw) purple!5Reference !5 !5 !5 black!90 AudioLDM [-1pt](ICML, 2023) black!20 T-FOLEY* [-1pt](ICASSP, 2024) black!20 ThinkSound [-1pt](NeurIPS, 2025) black!20 A2SB* [-1pt](ArXiv, 2025) black!20 AudioX [-1pt](ICLR, 2026) black!90 ∗ Generated clip duration is 4 seconds, following the model restrictions. Figure 7: Additional qualitative comparison of ATA variations using mel-spectrograms on three representative ESC-50 classes: crow, chirping birds, and hand_saw. Each row corresponds to one baseline and each column to one reference class. The figure highlights differences in local texture preservation, temporal organization, and variation behavior across methods. For A2SB inpainting, the masked and regenerated regions are outlined, since only a localized segment of the reference is modified. Black boxes indicate prominent texture regions in the reference, while white boxes highlight similar texture patterns preserved in the generated outputs. black Animals (crow) Natural Sounds (chirping birds) Exterior Sounds (handsaw) blue!5Reference !5 !5 !5