Paper deep dive
AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities
Sandy Abdo, Bill Kapralos, Priyamvada Tripathi, KC Collins, Adam Dubrowski
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability. Artificial intelligence (AI)-driven audio generative models are rapidly growing in popularity and have the potential to transform the way sound is synthesized and used across various applications. In response to this growing momentum, this chapter reviews and analyzes recent AI-based generative models for sound effect synthesis, with a focus on how different input modalities (text, visual, audio, and multimodal) affect the quality, controllability, and contextual relevance of the generated audio. It examines 30 peer-reviewed articles sourced from Google Scholar, IEEE Xplore, and the ACM Digital Library, exploring the evolution of AI generative models over the past five years. The results show that multiple models achieved state-of-the-art performance, producing high-fidelity, semantically aligned, and increasingly temporally coherent sound effects across tasks. However, despite these advances, the review identifies persistent challenges, including limitations in temporal synchronization for complex multi-event scenarios, gaps between objective metrics and human perception, and trade-offs between controllability and generative diversity. Overall, the chapter highlights that AI-driven sound effect generation is progressing toward more adaptive, scalable, and context-aware systems, offering significant implications for future sound design workflows and interactive media applications.
Tags
Links
- Source: https://arxiv.org/abs/2608.03742v1
- Canonical: https://arxiv.org/abs/2608.03742v1
Trouble viewing inline? Open PDF directly →
Full Text
88,825 characters extracted from source content.
Expand or collapse full text
AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities Sandy Abdo [0009−0006−5128−1625] , Bill Kapralos [0000−0003−0434−3847] , Priyamvada Tripathi [0009−0005−5070−7420] , KC Collins [0000−0003−0695−7228] , and Adam Dubrowski [0000−0002−2074−0933] Abstract Sound effects play a crucial role in conveying actions, events, and environ- mental cues across digital applications, often requiring a high degree of variation and contextual adaptability. Artificial intelligence (AI)-driven audio generative models are rapidly growing in popularity and have the potential to transform the way sound is synthesized and used across various applications. In response to this growing momen- tum, this chapter reviews and analyzes recent AI-based generative models for sound effect synthesis, with a focus on how different input modalities (text, visual, audio, and multimodal) affect the quality, controllability, and contextual relevance of the gener- ated audio. It examines 30 peer-reviewed articles sourced from Google Scholar, IEEE Xplore, and the ACM Digital Library, exploring the evolution of AI generative models over the past five years. The results show that multiple models achieved state-of-the-art performance, producing high-fidelity, semantically aligned, and increasingly tempo- rally coherent sound effects across tasks. However, despite these advances, the review identifies persistent challenges, including limitations in temporal synchronization for complex multi-event scenarios, gaps between objective metrics and human perception, and trade-offs between controllability and generative diversity. Overall, the chapter highlights that AI-driven sound effect generation is progressing toward more adaptive, scalable, and context-aware systems, offering significant implications for future sound design workflows and interactive media applications. Sandy Abdo Ontario Tech University, Oshawa, ON, Canada, e-mail: sandy.abdo@ontariotechu.net Bill Kapralos Ontario Tech University, Oshawa, ON, Canada, e-mail: Bill.Kapralos@ontariotechu.ca Priyamvada Tripathi Durham College, Oshawa, ON, Canada, e-mail: Priyamvada.Tripathi@durhamcollege.ca KC Collins Carleton University, Ottawa, ON, Canada e-mail: kccollins@cunet.carleton.ca Adam Dubrowski Ontario Tech University, Oshawa, ON, Canada, e-mail: adam.dubrowski@ontariotechu.net 1 arXiv:2608.03742v1 [cs.SD] 4 Aug 2026 2Abdo et al. Keywords Artificial intelligence· generative model· sound effect 1 Introduction Sound is fundamental to how humans perceive, interpret, and interact with digital sys- tems, and recent advances in artificial intelligence (AI) are transforming how audio is created, processed, and experienced. Rather than functioning as a static or purely reactive element, modern audio systems increasingly leverage AI to become adaptive, generative, and context-aware. This enables sound, including dialogue, sound effects, ambient noise, and music, to dynamically respond in real-time to user input, environ- mental conditions, and system states [6]. Among these, sound effects play a particularly critical role in conveying events, ac- tions, and spatial cues, making them essential in applications such as video games, film, virtual simulations, and user interfaces where immediate auditory feedback enhances realism and user engagement [6, 38]. Designing dynamic and responsive sound effects, however, remains a complex and resource-intensive task. Sound designers must work through extensive audio databases, extract specific sound clips, and meticulously mix and manipulate sounds to achieve the desired auditory experience [6]. However, digital applications including games and films may require thousands of sound effects to create experiences that resemble the real- world, especially since human pattern recognition capabilities make repeated sounds, such as footsteps, quickly noticeable and detrimental to immersion if not sufficiently varied [6]. This repetitive workflow, combined with the need for constant variation, can be tedious and limit creative exploration. Automating some of these tasks using AI could streamline the process, allowing designers to focus on experimenting with alternative soundscapes and enhancing the overall audio experience [12]. As AI workflows become more integrated into digital applications, sound effect generation is increasingly shifting toward procedural and adaptive approaches [40]. This makes it more difficult for sound designers to anticipate which sounds will be needed and under what conditions, as user interactions or system-driven events can produce unexpected and novel combinations of actions. Rare or unanticipated interactions may occur that were not considered during the design process, creating a demand for flexible and responsive audio systems. AI could synthesize plausible interaction sounds (e.g., action, background, and ambient) based on learned models of real-world acoustics [40]. In such systems, sound must adapt dynamically not only to interactions but also to continuously changing conditions. Acoustic properties such as reverberation, occlusion, and layering may need to update in real-time as scenes evolve [31]. Additionally, contextual parameters, such as user state, system variables, or desired emotional tone, can influence how sound effects are generated and modified. For example, a character’s footsteps could vary depending on factors like movement intensity, physical load, or condition, resulting in more nuanced and responsive auditory feedback. At the same time, procedural systems can generate vast numbers of variations in actions, objects, and events, making it impractical to design unique sound effects for every possibility [31]. AI-Based Sound Effect Generation3 AI offers a scalable solution by dynamically generating and adapting sounds to match an effectively unlimited range of scenarios. Beyond large-scale productions, many digital applications are developed by smaller studios, independent creators, and researchers who may lack the resources or expertise required to produce high-quality sound effects. AI-driven tools can help bridge this gap by lowering the cost and technical barriers to audio production, enabling more accessible creation of rich and context-aware sound [12]. This is particularly valuable in specialized or underrepresented domains, where tailored sound design is needed but traditional resources are limited, allowing a broader range of creators to develop high-quality auditory experiences. The purpose of this chapter is to provide an overview of audio generative AI, and to identify the available tools and models in the current literature that can generate audio from prompts such as text, video, images, or other audio. This review maps current knowledge, identifies trends, and guides future work in audio generative AI. 2 Research Strategy This chapter is a review that follows the structure of a narrative literature review, as outlined by Oxman et al. [34]. A narrative review serves to examine the literature comprehensively, offering a broad summary of the field. This approach is especially beneficial for readers new to the subject, providing them with valuable insights [16]. When crafting this chapter, we focused primarily on journal and conference papers, rather than for instance patents that may describe proprietary tools. The search was carried out using the Google Scholar, IEEE Xplore, and the ACM Digital Library databases between 6 February 2024, and 4 March 2025. We limited scope to these papers published within the past five years due to the rapidly changing nature of the field. Earlier models were included only for context and were not part of the search pool. IEEE Xplore, the ACM Digital Library, and Google Scholar were selected as the primary databases for this chapter due to their comprehensive coverage of research in computer science, engineering, and AI, fields most relevant to generative audio models. IEEE Xplore and ACM are leading publishers of conference proceedings and journals in AI and signal processing, providing access to cutting-edge research and original studies. Google Scholar complements these sources by offering broad, interdisciplinary coverage. Searches in other databases, such as Scopus and Web of Science, largely returned results already indexed in IEEE or ACM databases, indicating significant overlap and reinforcing the suitability of the chosen databases. Additionally, the review focused on original, peer-reviewed research articles appear- ing in refereed journals and conference proceedings, written in English and present models capable of generating sound effects. Models that primarily generate speech or music were excluded, as these areas have already been extensively studied and are covered by several recent review articles [4, 7, 20, 27, 32]. The research query was: (“Ar- tificial intelligence” OR AI OR “language model” OR LLM OR “Deep Learning” OR “neural network”) AND (“Text-to-Audio” OR “video-to-audio” OR “visual-to-audio” OR “audio-to-audio” OR “image-to-audio”) AND (“sound effect” OR “audio effect” 4Abdo et al. OR “game audio” OR “game sound”) AND (design* OR creat* OR generat* OR develop*). 2.1 Results The search initially yielded 204 articles. After removing duplicates and non–peer- reviewed articles [39], 60 abstracts were screened to assess relevance, specifically whether the study introduced an audio generative model. To ensure the relevance and quality of the reviewed literature, predefined inclusion and exclusion criteria were applied during both abstract screening and full-text review stages. Eligible studies were required to be peer-reviewed journal or conference papers published within the last five years, written in English, and presenting original research on artificial in- telligence–based generative models capable of producing sound effects. Studies were excluded if they were non–peer-reviewed (e.g., preprints, theses, patents, or technical reports), review articles, surveys, or editorials. Additionally, works focusing primarily on speech or music generation, lacking a clear generative modeling component, or not producing audio as a primary output were omitted. Out of the 60 abstracts screened, 43 articles met the inclusion criteria. Following a detailed full-text examination, 30 of these were deemed relevant to the review, with further exclusions applied to studies lacking sufficient methodological detail or proper evaluation to ensure analytical rigor. Although this is a narrative review, the search and screening process was documented transparently, and a summary of the information flow through the phases of the review is provided in the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) flow diagram [17] shown in Figure 1. The remainder of this review is organized around four key themes identified in the literature: (i) text-to-audio models, which take text as input; (i) visual-to-audio models, which take images or videos as input; (i) audio-to-audio models, which use audio as input; and (iv) multimodal audio generative models, which incorporate multiple input modalities while producing audio as output. Each of these themes is discussed in a dedicated subsection. Table 1 provides an overview of the models that will be discussed in this chapter. Table 1: Summary of audio-generative models. ModelArchitectureEvaluation Metrics Attributes Text-to-Audio AudioLDMLDM + CLAPFD, IS, KL, OVL, REL Efficient, scalable sound generation with high audio fidelity TangoLDM + Flan-T5 + VAE FAD, KL, FD, OVL, REL High audio quality with strong prompt relevance on small datasets Continued on next page AI-Based Sound Effect Generation5 ModelArchitectureEvaluation Metrics Attributes AudioLDM2LOA + LDM + AudioMAE CLAPScore, FAD, KL, OVL, REL Domain-agnostic model enabling semantic audio representation SonifyARPbD+ LLM+ audioLDM Preliminary usability evaluation Interactive AR-based audio generation AuffusionPixel VAE + LDM FD, FAD, KL, IS, CLAPScore, OVL, REL T2I method adaptation for TTA generation Tango 2Diffusion-DPO + CLAP FAD, KL, IS, CLAPScore, OVL, REL Temporal precision and flexible audio generation across applications SRC-gAudioU-Net + LDM + VAE + HiFi-GAN FD, KL, IS, FAD, CLAPScore, OGL, REL, AQ Multi-sampling rate training for improved audio quality Re-AudioLDMRetrieval- Augmented LDM IS, FAD, KL, CLAPScore Retrieval-guided generation aligned with real-world sounds Stable Audio Open DiT +MLPsFDopenl3 score, CLAPScore Open-weight model for scalable audio generation PicoAudioU-Net + LDMMOS, FAD, F1, 퐿 푓푟푒푞 1 Fine-grained temporal controllability AudioComposerLDM + hierarchical diffusion F1, ACC, MAE, MOS Hierarchical semantic modeling for improved audio quality Video-to-Audio V2RA-GANGANODG, OSG, architectural evaluation, user study Direct waveform synthesis via regression-based modeling FoleyGANBigGAN + visual action recognition network Retrieval Accuracy, IS, FID, NDB, user study Visual-guided audio synthesis using spectrogram generation FRIERENRFM + VAE + BigVGAN FD, IS, KL, KID, FAD, Acc, MIS Temporally aligned audio via ODE-based sampling Continued on next page 6Abdo et al. ModelArchitectureEvaluation Metrics Attributes AutoSFXSAM + spectrogram autoencoder + cross-attention missing, redundant, mismatched sounds, rhythm alignment, acoustic similarity, user study Pixel-level audiovisual feature-based sound generation MIMOSAMulti-step pipeline Friedman, Wilcoxon, user study Interactive spatial audio generation for content creators FoleyGenEnCodec + visual encoder + Transformer decoder FAD, KLD, IB, OVR, REL, alignment Cross-attention-based multimodal audio synthesis SonicVisionLMVLM + diffusion model CLAP-top, Onset Accuracy, AP, Time Accuracy, IoU, IS, MKL, FID Event-driven audio generation using vision-language modeling MaskVATTransformer- based + RVQ codec FDD, FDM, and FAD, semantic alignment, temporal synchronization Masked token-based audio generation with temporal alignment AV-LDMMLP + VAE + LDM FAD, AV-Sync, CLAPScore, user study Ambient-aware controllable audio generation Smooth-FoleyAuffusion + CLIP FAD, MKL, CLIP Score, user study Frame-level feature integration for improved alignment STA-V2ALDM TTAFD, FAD, KL, IS, PAM, CLAPScore, AV, A, OQ, AQ, SA, TA Multi-level feature refinement for semantic and temporal alignment Continued on next page AI-Based Sound Effect Generation7 ModelArchitectureEvaluation Metrics Attributes LoVAAutoregressive diffusion method FAD, IS, MKL, audio quality, semantic relevance, consistency, overall evaluation Long-form coherent and temporally consistent audio generation TA-V2Amulti-modal approach + LLM + LDM IS, FID, FAD, MKL, alignment, semantic, temporal alignment Text-guided video-to-audio synthesis for improved coherence Audio-to-Audio NoiseBandNetDDSP-derived filter-bank + MLP MRSTFT, FAD, amplitude randomization, loudness transfer, user-defined control Filter-bank-based synthesis for time-varying sound generation CSTs (2 tools)Mixed-Initiative Creative Interfaces + GAN User Study (TA)Interactive AI-assisted sound design for creative workflows Multimodal CoDiComposable LDM + Latent alignment FD, IS, RELUnified latent-space multimodal generation, including audio generation QueryMintAIGPT-3.5 Turbo + DALL-E-2 + Whisper v2 + TTS-1 TtPy, ASPy, HPy, ROUGE, user study Integrated multimodal pipeline for versatile content generation AmphionAudioLDM + PicoAudio FD, IS, KLOpen-source toolkit for text- and audio-driven generation VAMGGAN + Latent alignment IS, FID, classification accuracy, user study Bidirectional audio-visual generation with cross-modal alignment 8Abdo et al. Fig. 1 PRISMA flow diagram summarizing the review process. 2.2 Evaluation Metrics for Audio Generation Models This section summarizes the objective and subjective metrics that will be presented in this chapter for evaluating AI-based sound effect generation models. Objective met- rics provide quantitative measures of audio quality, diversity, and statistical similarity between generated and real-world data. Commonly used measures include distribution- based metrics such as Fr ́ echet Distance (FD/FAD), Fr ́ echet Inception Distance (FID), Kernel Inception Distance (KID), and KL Divergence (KL/KLD), as well as quality and diversity indicators including Inception Score (IS). Additionally, alignment and tempo- ral accuracy are evaluated using metrics such as CLAP Score, F1 Score, and AV-Sync. In contrast, subjective metrics capture human perception of generated audio, focusing on aspects such as naturalness, fidelity, and alignment with input prompts. These in- clude overall impression (OVL/OVR), mean opinion score (MOS), and audio quality AI-Based Sound Effect Generation9 (AQ/OQ), alongside alignment-focused measures such as audio-text relevance (REL), semantic alignment (SA), and temporal alignment (TA). Together, these metrics form a comprehensive evaluation framework that reflects both computational performance and perceptual quality (Table 2). Table 2 Evaluation metrics across models. Objective MetricAbbreviationWhat It MeasuresSubjective MetricAbbreviationWhat It Measures Fr ́ echet DistanceFD / FADAudio quality / realismOverall ImpressionOVL / OVR Perceived overall quality Inception ScoreISDiversity + qualityAudio-Text RelevanceRELPrompt alignment (human) KL DivergenceKL / KLDDistribution similarityMean Opinion Score MOSNaturalness rating CLAP ScoreCLAPScoreText-audio alignmentAudio QualityAQ / OQSound fidelity (human) F1 ScoreF1Temporal event accuracySemantic Alignment SAContent match (human) Fr ́ echet Inception Dist.FIDFeature-level qualityTemporal AlignmentTATiming accuracy (human) Kernel Inception Dist. KIDDistributional fidelity AV-SyncAV-SyncAudio-visual synchronization 3 Text-to-Audio Generation Text-to-Audio (TTA) is a growing field that recognizes, synthesizes, and generates audio based on given text input [27]. To generate the audio, the provided text can consist of content, attribute, and style. The type of audio generated includes music, speech, and sound effects. The research in this area mainly focuses on text-to-speech (TTS), particularly the quality, efficiency, and control of speech sound [20, 27]. Some progress has also been made in text-to-music (TTM) [4, 7, 32]. Unlike speech and music, sound effects are the least explored. In this chapter, we are focusing on research specifically examining sound effect generation, and therefore we do not describe speech or music generation here. Liu et al. [25] introduced AudioLDM, a generation model that produces high-quality sound from text descriptions using a latent diffusion model (LDM) architecture. Unlike previous approaches that generate audio in the waveform space (a computationally in- tensive process), AudioLDM works in the latent space (a lower-dimensional representa- tion of the data capturing essential features), enabling more efficient and scalable sound generation while preserving audio fidelity. The model leverages Contrastive Language- Audio Pretraining (CLAP), which provides a shared embedding space for both text and audio, allowing the model to understand and align textual prompts with audio features effectively. Additionally, its training in the latent space of the audio coder greatly re- duces computational cost without sacrificing perceptual quality. This also enables the model to generate audio that is longer and more diverse than previous methods. Audi- oLDM achieved state-of-the-art performance on several audio generation benchmarks, especially for text-to-audio generation tasks. The results show that AudioLDM models significantly outperform baseline TTA systems such as DiffSound, which can model high-dimensional audio signals, and AudioGen [23], an autoregressive model trained 10Abdo et al. on waveforms, across both objective measures such as Fr ́ echet distance (FD), inception score (IS) and Kullback–Leibler divergence (KL), and subjective evaluations, includ- ing overall impression (OVL), audio-text relevance (REL), with the best-performing variant, AudioLDM-L-Full, achieving the lowest FD and highest human-rated audio quality. These results indicate that AudioLDM generates audio that is more natural and better aligned with the textual prompts compared to baselines. The model is also capable of generating a wide variety of sounds, including environmental noises, music clips, and synthetic effects. Ghosal et al. [13] introduced Tango, a LDM that leverages a Large language model (LLM) to enhance representational capabilities, fine-tuning, and the learning of com- plex concepts. Tango uses Flan-T5, an LLM, to generate mel-spectrogram tokens that represent the frequency content of an audio signal over time. These tokens are then processed by a pre-trained audio variational autoencoder (VAE) to reconstruct the mel- spectrogram and ultimately produce audio via a vocoder. Trained on the AudioCaps dataset [21], consisting of 45,438 audio clips with their paired caption, Tango outper- formed previous state-of-the-art text-to-audio models including AudioLDM across both objective (Fr ́ echet Audio Distance (FAD), KL, FD) and subjective (OVL, REL) metrics, through evaluations by six participants who rated the audio quality and its relevance to the input text. Although trained on smaller datasets, the model achieved particularly strong performance due to its use of the Flan-T5 text encoder, demonstrating better audio quality, relevance to prompts, and sample efficiency compared to baselines. Liu et al. [26] introduced AudioLDM 2, a TTA generation tool capable of produc- ing audio, music, and intelligible speech without the domain-specific biases that often limit broader applications, especially in complex scenarios. To overcome this limita- tion, the authors propose the “language of audio” (LOA), a vector representation of an audio clip’s semantic information. Unlike existing methods, LOA can theoretically capture both fine-grained acoustic details and coarse-grained semantic content. In other words, it can represent both what is being said, and the identity of the sound produced. AudioLDM 2 leverages a GPT-2 language model to translate this information into an audio-masked autoencoder (AudioMAE), a decoder pre-trained on diverse audio con- tent. The processed information is then synthesized into audio using a LDM built on AudioMAE features. To evaluate its performance, the authors compared AudioLDM 2 against other audio-generation systems in TTA, TTM, and TTS tasks. The model achieving top scores across objective (CLAPScore, FAD, KL) and subjective metrics (OVL and REL), and significantly outperformed previous state-of-the-art models, in- cluding Tango, on AudioCaps and MusicCaps in TTA and TTM tasks, and achieved comparable performance in TTS. Su et al. [41] introduced SonifyAR, a context-aware audio generation tool designed to enhance sound effects in augmented reality (AR). The system leverages Programming by Demonstration (PbD) to detect potential sound-producing interactions in AR. These interactions are then analyzed by sound acquisition models and a LLM to generate a text-based description that includes details about the user, virtual object, and real- world environment. Next, the sound acquisition models and LLM identify the material and surface type, which are then used to create a text prompt for AudioLDM [25] to generate contextually appropriate sound effects. To evaluate the system, the authors conducted a preliminary user study with eight participants, who provided feedback on AI-Based Sound Effect Generation11 SonifyAR’s usability and effectiveness. Overall, participants responded positively, with many expressing a willingness to use the tool and acknowledged its potential to enhance AR immersion. Xue et al. [50] introduce Auffusion, a TTA model that integrates LDM from text- to-image (T2I) tasks to enhance cross-modal alignment. Unlike previous models that overlooked fine-grained performance, the authors incorporate pixel VAE, a technique commonly used in T2I to improve the quality of AI-generated images. Auffusion uses a text prompt to generate a latent representation, which is then reconstructed into an image by the VAE decoder. This image is subsequently denormalized into a mel-spectrogram and synthesized into audio. Their results show that Auffusion outperforms previous state-of-the-art models on the objective measures including FD, FAD, KL, IS, and CLAPScore, and generalizes well to the Clotho test set, a benchmark derived from the Clotho audio captioning dataset [8], despite being trained on a much smaller dataset. Additionally, subjective evaluations including OVL and REL show superior text-audio alignment. Majumder et al. [29] introduced Tango 2, a TTA generative model designed to syn- thesize high-quality sound using deep learning techniques. Using a diffusion model, the system is trained on extensive datasets to learn intricate patterns and structures of audio, enabling realistic and coherent sound generation. With features such as high-fidelity syn- thesis, scalability, and potential real-time processing, the model offers flexibility across various audio applications, including music composition, speech synthesis, and sound design. The model uses CLAP, a model that learns acoustics from natural language supervision, and temporal perturbation strategies, allowing it to enhance its ability to generate semantically accurate and temporally coherent audio. The results indicate that Tango 2 outperforms both its predecessor Tango and the baseline AudioLDM2 across a wide range of objective and subjective metrics (REL and OVL), showing notable improvements in audio quality (FAD, KL, IS), semantic alignment (CLAPScore), and temporal precision (Strongly Temporally-Aligned evaluation Metric (STEAM) [47]), especially with complex multi-event prompts. These gains are largely attributed to direct preference optimization (DPO) fine-tuning and event/temporal data augmentation strate- gies in the Audio-alpaca dataset, a dataset of prompts with preferred and less accurate audio outputs for training better TTA models (https://huggingface.co/datasets/declare- lab/audio-alpaca), which prove critical in enhancing performance and model preference alignment. These findings highlight the breakthrough potential of this approach, making it a powerful tool for various domains requiring high-quality audio generation. Li et al. [24] proposed a multi-sampling rate model called SRC-gAudio that uses sampling rate and text prompt as a condition to enhance audio generation quality. They mapped time step and sampling rate into a one-dimensional embedding, and along with text embedding they provided this conditioning information to a U-Net model, a neural network model used for image segmentation [37]. The LDM then guided by the conditioned information, restores the audio latent to mel-spectrogram and audio waveform using VAE and HiFi-Generative Adversarial Network (HiFi-GAN) to generate audio. The evaluation of SRC-gAudio used objective metrics, FD, KL, IS, FAD, and CLAPScore, showing that joint training across sampling rates, most notably at 32 kHz and 48 kHz, improves performance compared to training separate models. Subjective evaluation, based on ratings from eight participants using OGL, REL, and 12Abdo et al. audio quality (AQ), further confirms that pre-training on low-sampling-rate data boosts audio quality and that SRC-gAudio outperforms baselines such as AudioLDM2 in high-sampling-rate generation. Yuan et al. [52] present the Re-AudioLDM model designed to enhance audio gener- ation by leveraging retrieved audio samples as references. It operates by first retrieving relevant audio samples from a large-scale text-audio dataset based on the input text. These retrieved examples serve as references that guide the synthesis process, ensuring the generated audio aligns more closely with real-world sounds. The model is comprised of two main components: i) a retrieval module that identifies similar audio samples using semantic similarity techniques, and i) a generative model that synthesizes new audio based on the condition of the retrieved references. Experimental results demonstrate that Re-AudioLDM outperforms state-of-the-art models including AudioGen, AudioLDM, and Tango on all evaluation metrics (IS, FAD, KL, and CLAP score) when enhanced with retrieval information, especially excelling in generating semantically relevant and high-quality audio. It demonstrates strong performance on long-tailed and zero-shot audio generation tasks, outperforming traditional mixup strategies by effectively lever- aging retrieved audio-text pairs to improve robustness and generalization. Evans et al. [10] created a publicly available open-weight TTA called Stable Audio Open that was trained on Creative Commons (C) licensed audio. The LDM is a diffusion-transformer (DiT) [11], consisting of stacked blocks connected in series to attention layers and gated multi-layer perceptrons (MLPs), a type of neural network that consists of multiple interconnected nodes. To ensure no copyrighted content, the authors analysed the soundtracks before using them in their model training process. Results indicate that Stable Audio Open outperforms comparable baselines on the AudioCaps Dataset, achieving the best FDopenl3 score for AudioGen and AudioLDM2, indicating more realistic and plausible sound generation, while also scoring highest in CLAP score, reflecting strong alignment with text prompts. Xie et al. [48] introduced PicoAudio, a diffusion-based TTA model capable of precise temporal controllability in audio generation. The model employs a latent diffusion pro- cess, similar to text-to-image models, but adapted for generating high-fidelity audio with fine-grained alignment to text prompts. It leverages a U-Net architecture conditioned on textual embeddings to iteratively refine audio samples, ensuring both semantic accuracy and precise timing. The temporal conditioning mechanism allows it to generate audio segments that align exactly with specified time intervals. This approach improves upon prior TTA models, which often struggle to maintain synchronization between textual descriptions and generated sound. Particularly, the models was tested using subjective metrics, including mean opinion score (MOS) assesses audio quality (naturalness, dis- tortion, event accuracy) and temporal controllability (timestamp/frequency accuracy), with 10 evaluators rating five clips per model, and objective metrics (FAD for audio quality, segment F1 score for timestamp control, and a frequency error metric 퐿 푓푟푒푞 1 for frequency control). The results show that PicoAudio significantly outperforms baselines such as AudioLDM2 and Audit, achieving precise temporal control and superior audio quality across single and multi-event tasks. Wang et al. [43] introduced AudioComposer, a diffusion-based model designed for fine-grained audio generation using natural language descriptions. It integrates a AI-Based Sound Effect Generation13 transformer-based text encoder to process detailed semantic representations of sound attributes, such as pitch, rhythm, and texture, which guide the synthesis process. The model leverages a hierarchical approach, to generate audio using a fine-grained natural language description, allowing it to capture both high-level structures and intricate tem- poral details. It builds upon diffusion-based architectures, incorporating transformer- based text encoders to understand detailed descriptions, refine noisy signals and generate corresponding audio. AudioComposer was trained on a large-scale dataset with numer- ous audio samples, allowing it to generalize various sound types, including speech, music, and environmental noises. It significantly outperforms prior models in audio generation tasks involving timestamp, pitch, and energy control, across both objective metrics (퐹 1 -score was used for event accuracy (ACC) for pitch and energy categoriza- tion, and mean absolute error (MAE) between pitch and energy was also evaluated), and subjective (MOS) metrics. This shows that the generated audio exhibits greater coherence, fidelity, and adherence to textual descriptions. Ablation studies confirm that its performance benefits come from its design, including flow-based diffusion and the use of diverse training data. The combination of hierarchical modeling, a powerful text encoder, and diffusion-based synthesis results in superior fidelity, better alignment with textual descriptions, and more flexibility in generating complex soundscapes. In summary, recent advances in TTA generation, particularly for sound effects, demonstrate a rapid shift toward more efficient, controllable, and semantically aligned models (Figure 2). Diffusion-based approaches operating in latent spaces, such as AudioLDM and its successors, have significantly improved audio fidelity and scalabil- ity, while the integration of large language models and retrieval mechanisms has en- hanced the understanding of complex textual prompts and rare sound events. Emerging techniques further address key limitations by introducing temporal precision, multi- sampling strategies, and cross-modal learning, enabling more fine-grained and context- aware sound synthesis. Despite these advancements, sound effect generation remains less mature than speech and music synthesis, with ongoing challenges in realism, tem- poral alignment, and data efficiency. Overall, the reviewed work highlights both the substantial progress made and the promising future directions for developing robust, high-quality, and versatile sound effect generation systems. Fig. 2 Dominant architectures and synthesis pipelines for TTA generation (n represents the number of studies/works in this group). 14Abdo et al. 4 Visual-to-Audio Generation This section includes studies that use AI to generate audio from visuals such as images and videos. Video-to-audio (V2A) uses AI to recognize the content and style of a video to generate background music, spatial audio and sound effects [27]. Liu et al. [28] present an end-to-end deep learning approach for generating raw audio from visual content using a Generative Adversarial Network (GAN)-based framework called V2RA-GAN. Unlike conventional methods that rely on spectrograms or physics- based models, this approach formulates sound synthesis as a regression problem, di- rectly predicting synchronized raw audio from silent videos. The model includes a video encoder that extracts visual features, which are mapped to audio waveforms using a GAN. The system is fully trainable without additional inputs, improving scalability and reusability. The model was evaluated quantitatively using the Objective Difference Grade (ODG) for audio quality, and the Objective Similarity Grade (OSG) for audio similarity. Architectural evaluations indicated that filter layers and a combination of L1 and least-squares (LS) loss functions improved performance, while skip connections and Gaussian noise offered no benefits. Optimization schemes such as de-noising and peak-match improved clarity in specific scenarios, and user studies showed that V2RA- generated audio is generally high-quality and well-synchronized with visuals, often indistinguishable from real recordings. Overall, the results of the experiments demon- strated that V2RA-GAN produced high-quality, synchronized sounds in real-time, with practical applications in sound design and dubbing. Ghose & Prevost [14] enhanced Foley generation by introducing FoleyGAN, a visu- ally guided, class-conditioned model that uses a GAN designed to generate synchronous sound for silent videos. The model integrates a video action recognition network, which extracts temporal action information from video frames, and a sound generation net- work based on the BigGAN architecture. This architecture synthesizes high-resolution spectrograms conditioned on visual cues, which are then converted into audio using an inverse short-time Fourier transform (ISTFT). FoleyGAN was evaluated through qual- itative and quantitative experiments. Quantitative metrics, including Sound Retrieval Accuracy, IS, Fr ́ echet Inception Distance (FID), and Number of Statistically-Different Bins (NDB), show FoleyGAN with visual guidance outperforms baseline models in generation quality. Extensive ablation studies confirmed the effectiveness of temporal action information and the BigGAN architecture. Additionally, phase coherence analy- sis and a human survey demonstrated that generated sounds were generally perceived as high-quality and well-synchronized with visuals, with some variation across event classes. The model outperformed baseline models, achieving a sound retrieval accuracy of 76%, high audio-visual synchronicity (81% based on human surveys), and overall improved performance in comparison to other audio-visual synthesis methods. Wang et al. [44] introduced FRIEREN, a V2A model that uses rectified flow matching (RFM) to produce temporally aligned audio from silent video frames. It uses VAE to compress the mel-spectrogram in a continuous flow via vector field estimator, which is then converted to waveforms by the BigVGAN vocoder. RFM constructs straight- line flow trajectories, which allow for fast and accurate sampling when solved with an ordinary differential equation (ODE) solver. To allow for an audio-visual alignment, AI-Based Sound Effect Generation15 the vector field estimator uses a feed-forward transformer without temporal down- sampling, preserving the resolution of the temporal dimension. The model was evaluated quantitatively (FD, IS, KL, kernel inception distance (KID), FAD), and subjectively using MOS and outperformed prior models including Diff-Foley in inception score, audio fidelity, and alignment ACC, while being over nine times faster in one-step generation. Wang et al. [45] introduced AutoSFX, a tool designed to automate sound design for videos. AutoSFX consists of two main components: i) sound generation, i) and sound optimization. For sound generation, the tool follows the Segment Anything Model (SAM) [22], extracting compact, pixel-wise audiovisual features from the video. A spectrogram autoencoder then predicts and generates the corresponding audio based on the provided visual prompt. For sound optimization, AutoSFX employs a cross-attention mechanism to align auditory and visual information, ensuring better synchronization between generated sounds and video content. To evaluate AutoSFX’s performance, the authors conducted a quantitative study using metrics such as missing, redundant, and mismatched sounds, rhythm alignment, and acoustic similarity, and compared the results to state-of-the-art methods. Results showed that AutoSFX achieved the lowest error rate among the tested models on the VEGAS [55] and VGGSound datasets [3]. Additionally, AutoSFX outperformed the Diff-Foley model in generating animal sounds but was outperformed by Diff-Foley in generating instrumental sounds. The authors also conducted a user study with 30 participants, 10 video creators and 20 video viewers, who rated how well the generated soundtracks matched the visual content on a five- point scale. AutoSFX performed better than Pika, a free online audio generation tool, but fell short compared to professionally designed soundtracks. Despite this, 75% of participants perceived the generated sounds as realistic, 80% of viewers considered AutoSFX “very helpful” for amateur content creators, and 30% of video creators found it “very helpful” for automating complex sound design processes. Ning et al. [33] created an audio generation tool called Magnifying Immersion by Manipulating Objects in Spatial Audio (MIMOSA) to help amateur content creators generate spatial sounds (categorized as 3D localization or the creation of immersive experience) for videos. Their model uses a multi-step pipeline consisting of object de- tection, depth estimation, sound separation, audio tagging, and spatial audio rendering. This allows the users to interact with the tool, validate and correct errors in the AI- generated spatial audio by adjusting visual overlays in 2D and 3D manipulation panels. The model was evaluated by eight evaluators who rated five types of audios (Raw, Monaural, Default MIMOSA, and User-edited MIMOSA) on immersion and realism Statistical analysis using the Friedman test and Wilcoxon signed-rank tests. Results revealed that User-edited MIMOSA achieved significantly higher immersion than all other types, while default MIMOSA achieved immersion ratings comparable to raw audio. For realism, raw audio was rated highest, with computationally generated audio showing reduced realism due to perceived distortion. A user study was also conducted to examine how the tool assisted 15 novice content creators during their audio editing process. MIMOSA received high scores across metrics: usefulness, immersiveness, expressiveness, and ease of use. Participants praised the system’s real-time feedback, intuitive 2D/3D manipulation panels, and visual cues for error detection. While most fa- vored graphical manipulation, some preferred numerical inputs, indicating flexible user 16Abdo et al. needs. Suggestions included better object selection tools and an action history panel for improved editing workflow. Participants found the tool easy to use, supports cre- ativity, and improves immersion. They also concluded that allowing the users to refine the AI generated audio significantly enhances immersion while maintaining realism, highlighting the importance of human-AI collaboration. Mei et al. [30] introduced FoleyGen, an advanced V2A generation model designed to produce high-quality, visually aligned soundscapes. Built on a language modeling framework, it incorporates EnCodec, a neural audio codec for bidirectional waveform- token conversion, a visual encoder for feature extraction, and a Transformer-based audio decoder. Two variants, FoleyGen-C and FoleyGen-P, differ in their handling of visual features, FoleyGen-C employs a cross-attention mechanism, while FoleyGen- P integrates features via self-attention. FoleyGen was evaluated using both objective metrics, FAD, KLD, and ImageBind (IB) score, and subjective human evaluations assessing overall quality (OVR), REL, and alignment. FoleyGen outperformed prior models (SpecVQGAN and IM2WAV) across all metrics, achieving lower FAD and KLD scores and higher IB scores. Specifically, FoleyGen-C showed superior perfor- mance compared to FoleyGen-P, due to enhanced cross-attention mechanisms and better integration of visual cues. The best results were achieved using all-frame visual attention and multi-modal pretrained encoders such as CLIP and IB. Experiments with different visual encoders and attention mechanisms highlighted the advantages of multi-modal pretraining and all-frame attention in achieving better temporal alignment. Despite its success, challenges in perfect audio-video synchronization remain, suggesting direc- tions for future improvements. Xie et al. [49] proposed a novel framework, SonicVisionLM, designed to generate audio for silent videos using vision-language models (VLMs). Instead of directly syn- thesizing audio from visual data, the model first identifies relevant events in a video using a VLM, which then suggests appropriate sound effects. This approach simplifies the complex task of aligning video and audio by breaking it down into well-studied subproblems: image-to-text and text-to-audio mapping, using diffusion models. A key innovation is the introduction of a time-controlled audio adapter, ensuring better syn- chronization of generated sounds with video events. The SonicVisionLM model was evaluated using the Greatest Hits and CountixAV datasets across both conditional and unconditional sound generation tasks. For conditional generation tasks, it outperformed the prior state-of-the-art model on metrics such as CLAP-top, Onset Accuracy, Onset Average Precision (AP), Time Accuracy, and Intersection over Union (IoU). For uncon- ditional generation, the model surpassed existing video-to-audio generation models in IS and Mean KL (MKL), especially on the Greatest Hits dataset, using the IS, FID, MKL, and IoU metrics. Subjective evaluations further confirmed its superior performance in audio quality, temporal alignment, and synchronization. Pascual et al. [35] propose masked generative video-to-audio transformers (Mask- VAT), a model that generates audio from silent videos using a Transformer-based architecture that predicts tokenized audio sequences via masked generative token mod- eling. The model employs a Residual Vector Quantization (RVQ) codec, a hierarchi- cal tokenizer to compress the audio and embedded using pre-trained Descript Au- dio Codec (DAC) and return a codegram. The authors tested three variants of the model: i) 푀푎푠푘푉퐴푇 퐴푑푎퐿푁 , which uses AdaLN blocks, a modulator used in diffusion AI-Based Sound Effect Generation17 transformers, that matches the length of the transformer input using interpolation; i) 푀푎푠푘푉퐴푇 푆푒푞2푠푒푞 , a sequence-to-sequence model that uses a transformer encoder, bidi- rectional encoder representation from audio transformers (BEATs), for visual embed- dings and then uses cross-attention blocks, a parallel decoder, to mix the conditions with the main token, allowing auxiliary loss for synchronization; and i) 푀푎푠푘푉퐴푇 퐻푦푏푟푖푑 , combining both strategies where BEATS encoder to process visual features and AdaLN to align derived semantic features and S3D-derived alignment-sensitive features. The model was evaluated both objectively and subjectively on audio quality (FDD, FDM, and FAD), semantic alignment (novelty score and SparseSync), and temporal synchro- nization. Objective evaluation demonstrated that MaskVAT models, particularly the Seq2Seq and Hybrid variants, excelled in full-band audio quality. For semantic rele- vance, Seq2Seq and Hybrid variances resulted in the highest CLIP-Based similarity scores whereas alignment metrics resulted in higher scores for variants with AdaLN. Subjective evaluation confirmed these trends and indicated that 푀푎푠푘푉퐴푇 퐻푦푏푟푖푑 was rated highest in alignment and overall quality, and competitive with V2A-Mapper in fidelity and relevance, particularly among expert listeners. Overall, the results indi- cated that 푀푎푠푘푉퐴푇 퐻푦푏푟푖푑 as the most balanced and effective model for generating temporally aligned, semantically coherent, high-quality audio from video. Chen et al. [2] introduced AV-LDM, an ambient aware audio generative model that generates sounds from silent egocentric videos. The model introduces a training strategy that conditions audio generation on neighbouring audio clips to factor out ambient sounds. The audio waveform is first converted into a mel-spectrogram and then compressed into a latent representation using a VAE encoder. For conditioning, audio from a different timestamp (the audio condition) is processed using the same VAE encoder and further transformed into a fixed-length vector using a multilayer perceptron (MLP). Similarly, the input video is passed through a pre-trained video encoder to extract relevant visual features. These combined audio and video features are used to condition a latent diffusion model, enabling controllable generation of ambient sounds and enhancing the synthesis of action-focused audio. The model’s performance was evaluated both objectively, using FAD, Audio-visual synchronization (AV-Sync) and CLAP scores, and subjectively using human evaluation. Results showed that the model outperformed baselines such as Diff-Foley, Spec-VQGAN, REGNET, and retrieval-based models. Human evaluation with 20 participants further confirmed AV-LDM’s superior synthesis of semantically relevant, synchronized action sounds with controllable ambient noise, demonstrating promising generalization to VR game environments. Zhang et al. [54] proposed Smooth-Foley, a V2A generation model with improved semantic (alignment of generated sound with video content) and temporal (audio that is synchronized with the video) alignments. The authors established this by improving the temporal condition’s accuracy by using textual labels and by using a high-resolution frame-wise video embeddings instead of clip-wise. They integrated Auffusion, a TTA model with two lightweight adapters: i) a frame adapter, i) and a temporal adapter (Xue et al., 2024). The frame adapter improves semantic and temporal alignment by incor- porating high-resolution frame-wise video features instead of clip-wise embeddings. The temporal adapter refines synchronization by computing similarities between video frames and textual labels using CLIP, providing more accurate temporal conditions. 18Abdo et al. These adapters were trained separately while keeping the base TTA model frozen, al- lowing efficient adaptation. The model was evaluated used both objective and subjective evaluations. Objective measures included FAD, MKL, and CLIP Score. Subjectively, ten expert evaluators rated models on semantic alignment, temporal alignment, and audio quality. Smooth-Foley outperformed baseline models on both objective and sub- jective metrics across VGG and VGG-C datasets, with frame-wise features, yielding the best results in semantic alignment, audio fidelity, and temporal synchronization. Qual- itatively, the model captured complex audio events more accurately and demonstrated superior handling of temporal dynamics and realism. Ren et al. [36] addressed semantic and temporal alignment challenges with their STA-V2A model. They employed an LDM TTA framework and a cross-modal guidance approach that integrated both text and video to ensure proper alignment. To mitigate the interference of redundant information, they refined features at both local and global levels. For local refinement, they used a pre-text task that predicted pseudo-labels of audio onset from video, enabling the acquisition of localized video features. For global refinement, they extracted semantic features from videos using an attentive pooling module. The model’s performance was evaluated using both objective and subjective measures. Objective evaluations used the FD, FAD, KL, IS tools, prompting audio- language models (PAM), CLAP, audio-video alignment (AV-Align), and audio-audio alignment (A-Align). Results showed that STA-V2A achieves the lowest FAD and strong scores across the board, though it trails slightly behind FoleyCrafter, another V2A model, in PAM. Subjective evaluations by six expert annotators on quality (OQ), audio (AQ), semantic (SA), and temporal alignment (TA) confirmed that STA-V2A resulted in the highest scores and narrow 95% confidence intervals. Ablation studies further demonstrate that integrating filtered data, ControlNet with video features, onset- driven local features, and global semantic cues significantly enhances performance in audio quality, semantic alignment, and especially temporal synchronization. Overall, the STA-V2A model outperformed all baseline models across most objective and subjective metrics, including Diff-Foley. Cheng et al. [5] present Long-form Video-to-Audio (LoVA), a model designed for long-form video-to-audio generation. It integrates a video encoder with an au- toregressive audio generator to produce coherent and temporally aligned soundtracks for extended video sequences. The model is trained on diverse video-audio datasets, leveraging self-supervised learning techniques to enhance synchronization and realism. LoVA was tested using both objective metrics, FAD, IS, and MKL, and subjective human evaluations across four aspects (audio quality, semantic relevance, consistency, and overall). Results demonstrated that LoVA outperforms existing methods in gener- ating high-quality, context-aware audio, maintaining consistency over long durations and requiring the fewest inferences per audio. Experiments indicated its effectiveness in producing naturalistic sounds that closely match video content, setting a new bench- mark for long-form audio generation. Overall, LoVA achieved best overall performance, excelling in both quantitative and qualitative assessments. You et al. [51] designed TA-V2A, a textually assisted video-to-audio generation model, that creates realistic audio by leveraging both visual and textual inputs. It was developed using a multi-modal learning approach, integrating deep learning techniques to synthesize coherent soundscapes based on video content and supplementary textual AI-Based Sound Effect Generation19 descriptions. The model was evaluated using both objective and subjective assessments to benchmark its audio generation performance. Objective evaluation was conducted using four metrics, IS, FID, FAD, and MKL, on the VGGSound dataset, generating 8 s clips. Additionally, alignment accuracy was used to measure audio-video synchro- nization. TA-V2A, especially when using video-audio-language pretraining (CVALP) with concatenation, outperformed baseline models across all metrics, showing strong semantic fidelity and alignment, with particularly low FID and FAD values. In the sub- jective evaluation, 20 participants rated the audio quality using MOS for both semantic consistency and temporal alignment. TA-V2A with user-provided prompts achieved the highest scores, demonstrating that text-controlled inference significantly enhances perceived audio-video coherence. In summary, the reviewed studies demonstrate rapid progress in V2A generation, with modern approaches achieving increasingly high levels of audio quality, semantic relevance, and temporal synchronization (Figure 3). Techniques such as GANs, dif- fusion models, and transformer frameworks have each contributed unique strengths, from real-time waveform synthesis and high-fidelity spectrogram generation to im- proved alignment through cross-modal guidance and textual conditioning. Notably, recent work emphasizes not only automation but also controllability and human-AI collaboration, as seen in systems that incorporate user feedback or textual prompts to refine outputs. Despite these advances, challenges remain in achieving perfectly real- istic soundscapes and flawless synchronization across diverse and complex scenarios. Overall, the literature highlights a clear trajectory toward more robust, scalable, and user-adaptive V2A systems, with promising applications in content creation, virtual environments, and automated sound design. Fig. 3 Dominant architectures and synthesis pipelines for V2A generation (n represents the number of studies/works in this group). 5 Audio-to-Audio Generation Audio-to-audio generation in AI refers to systems that take an existing audio signal as input and produce a transformed or newly synthesized audio signal as output, rather than generating sound from text or images. It is often used to create variations of a reference sound or to apply stylistic or semantic changes while preserving key characteristics of the original audio [27]. Barahona-R ́ ıos & Collins [1] present NoiseBandNet, a novel neural audio synthesis method designed for generating time-varying sound effects. It builds on Differentiable Digital Signal Processing (DDSP) [9] (an architecture that allows the integration of a 20Abdo et al. signal processing component with deep learning techniques) but replaces the harmonic- plus-noise synthesizer with a filter-bank-based approach. This architecture processes white noise through a precomputed multi-band finite impulse response (FIR) filter- bank, generating M-band noise bands that span the entire frequency spectrum. The neural network predicts time-varying amplitude values for each band, conditioned on high-level audio features such as loudness and spectral centroid. During synthesis, these amplitudes are up-sampled and applied to the noise bands, which are then summed to reconstruct the target audio. This model is trained to predict time-varying amplitudes for multiple noise bands, which are then combined to reconstruct the target audio. The study evaluated NoiseBandNet’s performance against various DDSP configurations us- ing two objective metrics, i) Multiresolution Short-Time Fourier Transform (MRSTFT) loss, and i) FAD, and found that NoiseBandNet significantly outperforms all four tra- ditional DDSP noise synthesizers variants in MRSTFT loss. Although FAD followed similar trends, DDSP512 taps slightly outperformed NoiseBandNet on the pottery cate- gory, likely due to small dataset artifacts. Additional experiments, including amplitude randomization, loudness transfer and user-defined control parameters demonstrated creative sound design applications, highlighting NoiseBandNet’s superior performance and flexibility for expressive, controlled sound synthesis. Kamath et al. [19] created two AI-based Creative Support Tools (CSTs), which are AI models that allow sound designers to interact with and experiment with the StyleGAN model to assist with their creative practice. Their goal was not to compare these two in- terfaces, but rather, to provide expert sound designers with two AI tools to interact with. The first interface allows sound designers to generate a more realistic sound by inputting a synthetic sound designed by manipulating acoustic parameters (e.g., frequency and impulse width). The second interface uses technology-specific controls to generate the edited sound [46]. Using a qualitative approach, the authors conducted semi-structured interviews where they asked nine professional sound designers to provide their feedback on each of the CSTs. To analyze the interview transcripts, the authors used reflexive thematic analysis (TA), employing Atlas.TI (https://atlasti.com) for semantic coding and affinity diagramming to organize 76 codes into 12 themes. Rather than aiming for data saturation, the study prioritized participant diversity. The analysis revealed that the tool enables rapid iteration, generation of unique sounds, and minimizes reliance on manual recordings. While the tool’s unpredictability can spark creative ideas, it can also make it harder to complete tasks that require precision. Designers often prioritize artistic expression, even when the AI is focused on producing realistic sounds. As a result, although they value the creative possibilities the tool offers, they stress the im- portance of having more control and predictability, as well as tools that support, rather than replace, their creative decision-making. In summary, recent advancements in audio-to-audio generation demonstrate a clear shift toward more expressive, controllable, and application-oriented sound synthesis systems (Figure 4). Approaches such as NoiseBandNet highlight how integrating signal processing principles with deep learning can significantly improve synthesis quality and flexibility, particularly for complex, time-varying sounds. At the same time, the development of creative support tools underscores the importance of human-centered design, where AI serves as a collaborator rather than a replacement. Together, these works suggest that the future of audio-to-audio generation lies not only in improving AI-Based Sound Effect Generation21 technical performance, but also in enhancing user control, interpretability, and creative empowerment, enabling more seamless integration of AI into professional sound design workflows. Fig. 4 Dominant architectures and synthesis pipelines for audio-to-audio generation (n represents the number of studies/works in this group). 6 Multimodal Audio Generation This section discusses approaches that employ various forms of user input to generate various outputs including audio. Since our goal is audio generation tools, we focus on aspects of the model that generate audio output. Tang et al. [42] introduced Composable Diffusion (CoDi), a multimodal generative model that can generate any combination of output modalities (text, image, audio, and video) from any combination of input modalities. The model employs a LDM that is trained independently and then integrated through the ‘Bridging Alignment’ technique that employs text as an anchor modality to align conditional encoders. Additionally, the modalities are synchronized using a technique called latent alignment that allows a modality-specific environment encoder into a shared latent space of other modalities. The model was tested objectively under single and multi-condition generation. For single-condition generation, each modality was tested separately. Audio generation uses a VAE decoder and a vocoder to reconstruct audio samples from mel-spectrograms. In multi-condition generation, where the model generates two modalities simultaneously, it employs a technique called latent alignment. This allows each modality-specific encoder to project into a shared latent space, enabling cross-modal interaction. Additionally, a cross-attention layer is added to each modality’s U-Net, allowing each to access the latent features of the others. The model was evaluated quantitatively using metrics such as FD, IS, and Relevance (REL), demonstrating superior fidelity, relevance, and cross-modal understanding. Ghosh & Deepa [15] introduced QueryMintAI, a multimodal LLM designed to pro- cess various types of user inputs (text, images, videos, documents, URLs, audio, and databases) and generate corresponding outputs across different modalities. It leverages OpenAI’s GPT-3.5 Turbo (text-to-text generator), DALL-E-2 (text-to-image generator), Whisper v2 (speech-to-text), and TTS-1 (text-to-speech generator), among others, to fa- cilitate seamless interactions, allowing it to generate various forms of outputs. Through the model, the authors aimed to enhance user experience by providing a unified, intelli- gent, and private AI assistant capable of handling diverse data formats. They compared the model to other leading AI systems, including ChatGPT (GPT-3.5), Gemini, CoPi- 22Abdo et al. lot, and ClaudeAI. The results showed that QueryMintAI outperformed these models in both multimodal capabilities and overall performance, as measured by key metrics such as Total Text Perplexity (TtPy), Average Sentence Perplexity (ASPy), Highest Perplex- ity (HPy), and ROUGE scores (i.e., ROUGE-1 (R1), ROUGE-2 (R2), and ROUGE-L (RL)). User evaluation with 50 participants assessed the relevance, fluency, coherence, and overall quality of the model, which were aligned with the objective results. Overall, the study concluded that QueryMintAI is a highly versatile, multimodal AI model that delivers improved performance across multiple domains, making it a powerful tool for personal AI applications, accessibility, and creative content generation. Zhang et al. [53] introduced Amphion, a user-friendly, open-source toolkit designed for generating audio, speech, and music. Amphion aims to help both beginners and researchers explore generative AI with ease. It supports multiple tasks, including TTS, TTM, TTA, and audio-to-audio generation. We focus on TTA and audio-to-audio gen- eration. For TTA, Amphion uses pre-trained models like AudioLDM and PicoAudio to generate audio from descriptive text. The model achieved better performance on FD, IS, and KL scores in comparison to open-source models such as Diffsound and Audi- oLDM. The model also supports audio-to-audio generation, referred to as wavelength- to-wavelength, processes auditory input to perform tasks such as voice conversion, singing voice conversion, emotion conversion, accent conversion, and speech transla- tion. In their paper, the authors only evaluated singing voice conversion where they had the model convert 48 singing utterances into a female and a male target. Results showed that the model outperformed SoftVC (https://github.com/bshall/soft-vc) in both naturalness and speaking similarity. To evaluate its performance, Amphion was com- pared against other open-source models and found that it consistently delivered superior results. Hao et al. [18] present the visual-audio mutual generation (VAMG), dynamic cross- modal model designed to compensate for missing audio or visual information in videos. The model facilitates both audio-to-visual and visual-to-audio generation while also al- lowing self-generation within each modality. It uses GANs with joint optimization of modal reconstruction and adversarial constraints to address issues of structural alignment and signal compensation. The trained model demonstrated effectiveness in instrument- and pose-oriented audio-visual generation. VAMG significantly outper- formed prior methods for generating images from corresponding sounds (S2IC) and for generating sounds from corresponding images (I2S) in generation quality, as quantified by IS, FID, and classification accuracy. Additionally, subjective evaluation by 21 par- ticipants showed significant improvements in the satisfactory and acceptable ratings of the generated modality compared to baseline. Through latent embeddings and a novel adversarial loss function, VAMG enables dynamic modality compensation, ensuring more accurate and coherent cross-modal translations. In summary, these models highlight a clear trend toward increasingly flexible and uni- fied multimodal systems that incorporate diverse user inputs to produce high-quality au- dio outputs (Figure 5). Models such as CoDi emphasize tightly coupled latent represen- tations for synchronized cross-modal generation, while frameworks like QueryMintAI demonstrate the effectiveness of integrating specialized models into a cohesive pipeline for practical applications. Toolkits such as Amphion further lower the barrier to en- try by providing accessible implementations of advanced audio generation techniques, AI-Based Sound Effect Generation23 and approaches like VAMG push the boundaries of cross-modal reasoning by enabling bidirectional generation and compensation between modalities. Collectively, these mod- els illustrate that modern audio generation systems are moving beyond isolated tasks toward holistic, multimodal frameworks that leverage shared representations, cross- modal alignment, and user-centric design to achieve more coherent, context-aware, and versatile audio synthesis. Fig. 5 Dominant architectures and synthesis pipelines for multimodal generation (n represents the number of studies/works in this group). 7 Discussion This chapter highlights the rapid advancement of AI-driven generative models for sound effect synthesis, particularly over the past five years. Across text-to-audio, visual- to-audio, audio-to-audio, and multimodal systems, a clear trend emerges toward in- creasingly sophisticated models that can generate high-quality, contextually relevant, and temporally aligned audio (Table 3). Diffusion-based architectures, especially latent diffusion models (LDMs), consistently demonstrate strong performance across mul- tiple modalities, largely due to their ability to balance computational efficiency with high-fidelity output [25, 26]. One of the most significant findings is the improvement in semantic alignment between input prompts and generated audio. Models such as AudioLDM2, Tango 2, and PicoAudio demonstrate that integrating language models and multimodal embeddings enhances the system’s understanding of contextual cues, resulting in more accurate and expressive sound generation [26, 29, 48]. Similarly, in visual-to-audio tasks, models like FoleyGen and Smooth-Foley show that incorporating fine-grained temporal and frame-level features can significantly improve synchronization between visual events and generated sound [30, 54]. Despite these advances, several challenges remain. First, while objective metrics (e.g., FAD, IS, KL divergence) indicate strong performance, subjective evaluations often reveal gaps in perceived realism and quality [13,30,35]. This discrepancy suggests that current evaluation methods may not fully capture human auditory perception, emphasizing the need for more robust and perceptually grounded metrics. Second, temporal alignment, particularly in complex, multi-event scenarios, remains a persistent challenge, even for state-of-the-art systems [36, 48]. Another key limitation is the trade-off between controllability and creativity. While some models offer precise control over attributes such as timing, pitch, and energy, others prioritize generative diversity at the expense of predictability [1]. As highlighted 24Abdo et al. in the discussion of audio-to-audio systems and creative support tools, users, especially professional sound designers, value systems that enhance creativity while maintaining a degree of control and reliability [19]. Furthermore, accessibility and scalability are important considerations. AI-driven tools have the potential to democratize sound design by reducing the need for large audio libraries and specialized expertise [12]. However, many high-performing models still require substantial computational resources and large training datasets, which may limit their adoption in smaller studios or independent projects [10]. Finally, multimodal systems represent a promising direction for future research. By integrating multiple input types (e.g., text, video, and audio), these models can generate more coherent and context-aware outputs [15, 18, 42]. However, ensuring consistent alignment across modalities remains a complex challenge that requires further investigation. Table 3 Model distribution by category & subcategory. CategorySubcategory# ModelsDominant Architecture Text-to-AudioLatent Diffusion Models6LDM + language encoders Text-to-AudioSpecialized Control3LDM + temporal/attribute control Text-to-AudioApplication-Specific2LDM + domain adaptation Visual-to-AudioGAN-Based2GAN + visual encoders Visual-to-AudioDiffusion/Flow-Based4LDM/RFM + alignment modules Visual-to-AudioPipeline/Hybrid7Multi-component pipelines Audio-to-AudioSignal Processing1DDSP-derived filter-bank Audio-to-AudioCreative Tools1GAN + human-in-the-loop MultimodalCross-Modal / Toolkits4LDM + latent alignment 8 Conclusion This chapter examined recent developments in AI-based generative models for sound effect synthesis, focusing on text-to-audio, visual-to-audio, audio-to-audio, and mul- timodal approaches. The findings demonstrate that AI has significantly advanced the capability to generate realistic, diverse, and context-aware sound effects, offering trans- formative potential for applications such as video games, film production, virtual reality, and interactive systems. These advancements imply a fundamental shift in sound effect synthesis workflows, where generative AI can augment or partially replace traditional manual sound design, enabling rapid prototyping, iterative creativity, and on-demand audio generation tailored to specific scenes or user interactions. Diffusion-based models and multimodal learning frameworks have emerged as domi- nant approaches, enabling improved audio quality, semantic relevance, and adaptability. These technologies reduce the reliance on manual sound design processes and large cu- rated libraries, making high-quality audio production more efficient and accessible. As a result, sound effect creation is becoming increasingly democratized, allowing smaller AI-Based Sound Effect Generation25 studios and individual creators to produce professional-grade audio content without extensive resources, while also enabling adaptive and personalized soundscapes in real-time applications. However, challenges remain in achieving perfect temporal synchronization, improv- ing perceptual realism, and balancing user control with generative flexibility. Addition- ally, the gap between objective evaluation metrics and subjective human perception highlights the need for more comprehensive assessment frameworks. This underscores the continued importance of human oversight, creative direction, and hybrid workflows that combine AI generation with expert refinement to ensure artistic intent and narrative coherence are preserved. Overall, AI-driven audio generation represents a rapidly evolving field with signif- icant implications for the future of sound design. As these technologies continue to mature, they are likely to play a central role in shaping next-generation digital expe- riences, enabling richer, more immersive, and more dynamic auditory environments. This evolution points toward a future where sound design becomes more interactive, context-aware, and scalable, fundamentally redefining the role of sound designers from content creators to creative directors of generative systems. References 1. Barahona-R ́ ıos, A., Collins, T.: NoiseBandNet: controllable time-varying neural synthesis of sound effects using filterbanks. IEEE/ACM Transactions on Audio, Speech, and Language Processing 32, 1573–1585 (2024). DOI 10.1109/TASLP.2024.3364616. URL https://ieeexplore.ieee. org/document/10440034/ 2. Chen, C., Peng, P., Baid, A., Xue, Z., Hsu, W.N., Harwath, D., Grauman, K.: Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos. Computer Vision – ECCV 2024, Springer 15128, 277–295 (2025). DOI 10.1007/978-3-031-72897-6 16. URL https://link.springer.com/10.1007/978-3-031-72897-6_16 3. Chen, H., Xie, W., Vedaldi, A., Zisserman, A.: VGGSound: A Large-scale Audio-Visual Dataset (2020). DOI 10.48550/arXiv.2004.14368. URL http://arxiv.org/abs/2004.14368. ArXiv:2004.14368 [cs] 4. Chen, K., Wu, Y., Liu, H., Nezhurina, M., Berg-Kirkpatrick, T., Dubnov, S.: MusicLDM: En- hancing Novelty in text-to-music Generation Using Beat-Synchronous mixup Strategies. In: 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1206– 1210. Seoul, Republic of South Korea (2024). DOI 10.1109/ICASSP48485.2024.10447265. URL https://ieeexplore.ieee.org/document/10447265/ 5. Cheng, X., Wang, X., Wu, Y., Wang, Y., Song, R.: LoVA: Long-form video-to-audio generation. In: 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1–5. Hyderabad, India (2025). DOI 10.1109/ICASSP49660.2025.10888085. URL https: //ieeexplore.ieee.org/document/10888085/ 6. Collins, K.: Game Sound: An Introduction to the History, Theory, and Practice of Video Game Music and Sound Design.The MIT Press, Cambridge, Mass (2008). DOI 10.7551/mitpress/7909.001.0001. URL https://direct.mit.edu/books/book/2460/ Game-SoundAn-Introduction-to-the-History-Theory 7. Dash, A., Agres, K.: AI-based affective music generation systems: a review of methods and challenges. ACM Computing Surveys 56(11), 1–34 (2024). DOI 10.1145/3672554. URL https: //dl.acm.org/doi/10.1145/3672554 8. Drossos, K., Lipping, S., Virtanen, T.: Clotho: an Audio Captioning Dataset. In: ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 26Abdo et al. p. 736–740. IEEE, Barcelona, Spain (2020). DOI 10.1109/ICASSP40776.2020.9052990. URL https://ieeexplore.ieee.org/document/9052990/ 9. Engel, J., Hantrakul, L., Gu, C., Roberts, A.: DDSP: differentiable digital signal processing. CA, USA (2020). URL https://openreview.net/pdf?id=B1x1ma4tDr 10. Evans, Z., Parker, J.D., Carr, C., Zukowski, Z., Taylor, J., Pons, J.: Stable Audio Open. In: 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1–5. Hyderabad, India (2025). DOI 10.1109/ICASSP49660.2025.10888461. URL https: //ieeexplore.ieee.org/document/10888461/ 11. Evans, Z., Parker, J.D., Carr, C.J., Zukowski, Z., Taylor, J., Pons, J.: Long-form music generation with latent diffusion (2024). DOI 10.48550/arXiv.2404.10301. URL http://arxiv.org/abs/ 2404.10301. ArXiv:2404.10301 [cs] 12. Filimowicz, M.: Doing research in sound design, 1 edn. Focal Press, London (2021). DOI 10.4324/9780429356360. URL https://w.taylorfrancis.com/books/9780429356360 13. Ghosal, D., Majumder, N., Mehrish, A., Poria, S.: Text-to-audio generation using instruction guided latent diffusion model. In: Proceedings of the 31st ACM International Conference on Multimedia, p. 3590–3598. ACM, Ottawa ON Canada (2023). DOI 10.1145/3581783.3612348. URL https://dl.acm.org/doi/10.1145/3581783.3612348 14. Ghose, S., Prevost, J.J.: FoleyGAN: Visually Guided Generative Adversarial Network-Based Synchronous Sound Generation in Silent Videos. IEEE Transactions on Multimedia 25, 4508– 4519 (2023). DOI 10.1109/TMM.2022.3177894. URL https://ieeexplore.ieee.org/ document/9782577/ 15. Ghosh, A., Deepa, K.: QueryMintAI: Multipurpose multimodal large language models for personal data. IEEE Access 12, 144631–144651 (2024). DOI 10.1109/ACCESS.2024.3468996. URL https://ieeexplore.ieee.org/document/10695061/ 16. Grant, M.J., Booth, A.: A typology of reviews: an analysis of 14 review types and associ- ated methodologies. Health Information & Libraries Journal 26(2), 91–108 (2009). DOI 10.1111/j.1471-1842.2009.00848.x.URL https://onlinelibrary.wiley.com/doi/10. 1111/j.1471-1842.2009.00848.x 17. Haddaway, N.R., Page, M.J., Pritchard, C.C., McGuinness, L.A.: PRISMA2020: An R package and Shiny app for producing PRISMA 2020-compliant flow diagrams, with interactivity for optimised digital transparency and Open Synthesis. Campbell Systematic Reviews 18(2), e1230 (2022). DOI 10.1002/cl2.1230. URL https://onlinelibrary.wiley.com/doi/10.1002/cl2.1230 18. Hao, W., Guan, H., Zhang, Z.: VAG: A Uniform Model for Cross-Modal Visual-Audio Mu- tual Generation. IEEE Transactions on Neural Networks and Learning Systems 36(3), 4196– 4208 (2025). DOI 10.1109/TNNLS.2022.3161314. URL https://ieeexplore.ieee.org/ document/9753685/ 19. Kamath, P., Morreale, F., Bagaskara, P.L., Wei, Y., Nanayakkara, S.: Sound designer-generative AI interactions: towards designing creative support tools for professional sound designers. In: Proceedings of the CHI Conference on Human Factors in Computing Systems, p. 1–17. ACM, Honolulu HI USA (2024). DOI 10.1145/3613904.3642040. URL https://dl.acm.org/doi/ 10.1145/3613904.3642040 20. Kaur, N., Singh, P.: Conventional and contemporary approaches used in text to speech syn- thesis: a review.Artificial Intelligence Review 56(7), 5837–5880 (2023).DOI 10.1007/ s10462-022-10315-0. URL https://link.springer.com/10.1007/s10462-022-10315-0 21. Kim, C.D., Kim, B., Lee, H., Kim, G.: AudioCaps: generating captions for audios in the wild. In: Proceedings of the 2019 Conference of the North, p. 119–132. Association for Com- putational Linguistics, Minneapolis, Minnesota (2019). DOI 10.18653/v1/N19-1011. URL http://aclweb.org/anthology/N19-1011 22. Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W.Y., Doll ́ ar, P., Girshick, R.: Segment Anything (2023). DOI 10.48550/arXiv. 2304.02643. URL http://arxiv.org/abs/2304.02643. ArXiv:2304.02643 [cs] 23. Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., D ́ efossez, A., Copet, J., Parikh, D., Taigman, Y., Adi, Y.: AudioGen: Textually Guided Audio Generation (2023). DOI 10.48550/arXiv.2209.15352. URL http://arxiv.org/abs/2209.15352. ArXiv:2209.15352 [cs] AI-Based Sound Effect Generation27 24. Li, C., Xu, M., Yu, D.: SRC-gAudio: Sampling-Rate-Controlled Audio Generation. In: 2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), p. 1–6. IEEE, Macau, Macao (2024). DOI 10.1109/APSIPAASC63619.2025.10849319. URL https://ieeexplore.ieee.org/document/10849319/ 25. Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D., Wang, W., Plumbley, M.D.: AudioLDM: Text-to-Audio Generation with Latent Diffusion Models (2023). DOI 10.48550/arXiv.2301.12503. URL http://arxiv.org/abs/2301.12503. ArXiv:2301.12503 [cs] 26. Liu, H., Yuan, Y., Liu, X., Mei, X., Kong, Q., Tian, Q., Wang, Y., Wang, W., Wang, Y., Plumb- ley, M.D.: AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised Pretraining. IEEE/ACM Transactions on Audio, Speech, and Language Processing 32, 2871–2883 (2024). DOI 10.1109/taslp.2024.3399607. URL https://ieeexplore.ieee.org/document/10530074/. Publisher: Institute of Electrical and Electronics Engineers 27. Liu, Q., Chang, C., Shen, H., Cheng, S., Li, X., Zheng, R.: Research on artificial intelligence gener- ated audio. In: Sixth International Conference on Computer Information Science and Application Technology (CISAT 2023), vol. 12800, p. 1206–1212. SPIE (2023) 28. Liu, S., Li, S., Cheng, H.: Towards an End-to-End Visual-to-Raw-Audio Generation With GAN.IEEE Transactions on Circuits and Systems for Video Technology 32(3), 1299– 1312 (2022). DOI 10.1109/TCSVT.2021.3079897. URL https://ieeexplore.ieee.org/ document/9430540/ 29. Majumder, N., Hung, C.Y., Ghosal, D., Hsu, W.N., Mihalcea, R., Poria, S.: Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization. In: Proceed- ings of the 32nd ACM International Conference on Multimedia, p. 564–572. ACM, Melbourne VIC Australia (2024). DOI 10.1145/3664647.3681688. URL https://dl.acm.org/doi/10. 1145/3664647.3681688 30. Mei, X., Nagaraja, V., Le Lan, G., Ni, Z., Chang, E., Shi, Y., Chandra, V.: Foleygen: visually-guided audio generation. In: 2024 IEEE 34th International Workshop on Machine Learning for Signal Processing (MLSP), p. 1–6. IEEE, London, United Kingdom (2024). DOI 10.1109/MLSP58920. 2024.10734721. URL https://ieeexplore.ieee.org/document/10734721/ 31. Menexopoulos, D., et al.: The state of the art in procedural audio. Journal of the Audio Engineering Society 71, 826–848 (2023). DOI 10.17743/jaes.2022.0108 32. Mitra, R., Zualkernan, I.: Music Generation Using Deep Learning and Generative AI: A Systematic Review. IEEE Access 13, 18079–18106 (2025). DOI 10.1109/ACCESS.2025.3531798. URL https://ieeexplore.ieee.org/document/10845168/ 33. Ning, Z., Zhang, Z., Ban, J., Jiang, K., Gan, R., Tian, Y., Li, T.J.J.: MIMOSA: Human-AI co-creation of computational spatial audio effects on videos. In: Creativity and Cognition, p. 156–169 (2024). DOI 10.1145/3635636.3656189. URL http://arxiv.org/abs/2404.15107. ArXiv:2404.15107 [cs] 34. Oxman, A.D.: Users’ guides to the medical literature: Vi. how to use an overview. JAMA 272(17), 1367 (1994). DOI 10.1001/jama.1994.03520170077040. URL http://jama.jamanetwork. com/article.aspx?doi=10.1001/jama.1994.03520170077040 35. Pascual, S., Yeh, C., Tsiamas, I., Serr ` a, J.: Masked generative Vvideo-to-audio transformers with enhanced synchronicity. In: A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, G. Varol (eds.) Computer Vision – ECCV 2024, vol. 15145, p. 247–264. Springer Nature Switzerland, Cham (2025). DOI 10.1007/978-3-031-73021-4 15. URL https://link.springer.com/10. 1007/978-3-031-73021-4_15. Series Title: Lecture Notes in Computer Science 36. Ren, Y., Li, C., Xu, M., Liang, W., Gu, Y., Chen, R., Yu, D.: STA-V2A: Video-to-audio generation with semantic and temporal alignment. In: 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1–5 (2025). DOI 10.1109/ICASSP49660.2025. 10890132. ISSN: 2379-190X 37. Ronneberger, O., Fischer, P., Brox, T.: U-Net: Convolutional Networks for Biomedical Im- age Segmentation. In: N. Navab, J. Hornegger, W.M. Wells, A.F. Frangi (eds.) Medical Im- age Computing and Computer-Assisted Intervention – MICCAI 2015, vol. 9351, p. 234– 241. Springer International Publishing, Cham (2015). DOI 10.1007/978-3-319-24574-4 28. URL http://link.springer.com/10.1007/978-3-319-24574-4_28. Series Title: Lec- ture Notes in Computer Science 28Abdo et al. 38. Serafin, S., Franinovi ́ c, K., Hermann, T., Lemaitre, G., Rinott, M., Rocchesso, D.: Sonic Interaction Design, The sonification handbook, vol. 5. Logos Publishing House, Berlin (2011). URL https: //sonification.de/handbook/download/TheSonificationHandbook-chapter5.pdf 39. Sheppard,V.:Acceptablesourcesforliteraturereviews.In:Research MethodsfortheSocialSciences:AnIntroduction.Pressbooks(2020). URL https://pressbooks.bccampus.ca/jibcresearchmethods/chapter/ 5-3-acceptable-sources-for-literature-reviews/ 40. Steinmetz, C., Mitcheltree, C., Wichern, G., et al.: Audio signal processing in the artifi- cial intelligence era. Journal of the Audio Engineering Society 73, 406–428 (2025). DOI 10.17743/jaes.2022.0209 41. Su, X., Koh, E., Xiao, C.: Sonifyar: context-aware sound effect generation in augmented reality. In: Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, CHI EA ’24. Association for Computing Machinery, New York, NY, USA (2024). DOI 10.1145/3613905. 3650927. URL https://doi.org/10.1145/3613905.3650927 42. Tang, Z., Yang, Z., Zhu, C., Zeng, M., Bansal, M.: Any-to-any generation via composable diffusion. Advances in Neural Information Processing Systems 36, 16083–16099 (2023) 43. Wang, Y., Chen, H., Yang, D., Wu, Z., Wu, X.: AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions. In: 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1–5. Hyderabad, India (2025). DOI 10.1109/ICASSP49660.2025.10888303. URL https://ieeexplore.ieee.org/document/ 10888303/ 44. Wang, Y., Guo, W., Huang, R., Huang, J., Wang, Z., You, F., Li, R., Zhao, Z.: Frieren: Efficient video-to-audio generation network with rectified flow matching. Advances in Neural Information Processing Systems 37, 128118–128138 (2024) 45. Wang, Y., Wang, Z., Huang, H.: AutoSFX: Automatic Sound Effect Generation for Videos.In: Proceedings of the 32nd ACM International Conference on Mul- timedia, p. 9923 – 9932 (2024).DOI 10.1145/3664647.3681109.URL https: //w.scopus.com/inward/record.uri?eid=2-s2.0-85209818752&doi=10.1145% 2f3664647.3681109&partnerID=40&md5=e65175b6563a84f836716df38506b5f0.Type: Conference paper 46. Weisz, J.D., Muller, M., He, J., Houde, S.: Toward general design principles for generative AI ap- plications (2023). DOI 10.48550/ARXIV.2301.05578. URL https://arxiv.org/abs/2301. 05578. Version Number: 1 47. Xie, Z., Xu, X., Wu, Z., Wu, M.: AudioTime: A Temporally-aligned Audio-text Benchmark Dataset (2024). DOI 10.48550/arXiv.2407.02857. URL http://arxiv.org/abs/2407.02857. ArXiv:2407.02857 [cs] 48. Xie, Z., Xu, X., Wu, Z., Wu, M.: PicoAudio: Enabling Precise Temporal Controllability in Text- to-Audio Generation. In: 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1–5. Hyderabad, India (2025). DOI 10.1109/ICASSP49660.2025. 10890827. URL https://ieeexplore.ieee.org/document/10890827/ 49. Xie, Z., Yu, S., He, Q., Li, M.: Sonic VisionLM: playing sound with vision language models. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 26856– 26865. IEEE, Seattle, WA, USA (2024). DOI 10.1109/CVPR52733.2024.02537. URL https: //ieeexplore.ieee.org/document/10655167/ 50. Xue, J., Deng, Y., Gao, Y., Li, Y.: Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing 32, 4700–4712 (2024). DOI 10.1109/TASLP.2024.3485485 51. You, Y., Wu, X., Qu, T.: TA-V2A: Textually Assisted Video-to-Audio Generation. In: 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1–5. Hyderabad, India (2025). DOI 10.1109/ICASSP49660.2025.10887573. URL https: //ieeexplore.ieee.org/document/10887573/ 52. Yuan, Y., Liu, H., Liu, X., Huang, Q., Plumbley, M.D., Wang, W.: Retrieval-Augmented Text-to- Audio Generation. In: ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 581–585. IEEE, Seoul, Korea, Republic of (2024). DOI 10.1109/ICASSP48485.2024.10447898. URL https://ieeexplore.ieee.org/document/ 10447898/ AI-Based Sound Effect Generation29 53. Zhang, X., Xue, L., Gu, Y., Wang, Y., Li, J., He, H., Wang, C., Liu, S., Chen, X., Zhang, J., Fang, Z., Chen, H., Tang, T.Y., Zou, L., Wang, M., Han, J., Chen, K., Li, H., Wu, Z.: Amphion: an open- source audio, music, and speech generation toolkit. In: 2024 IEEE Spoken Language Technology Workshop (SLT), p. 879–884. Macao (2024). DOI 10.1109/SLT61566.2024.10832255. URL https://ieeexplore.ieee.org/document/10832255/ 54. Zhang, Y., Xu, X., Wu, M.: Smooth-Foley: Creating continuous sound for video-to-audio gener- ation under semantic guidance. In: 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1–5 (2025). DOI 10.1109/ICASSP49660.2025.10890403. ISSN: 2379-190X 55. Zhou, Y., Wang, Z., Fang, C., Bui, T., Berg, T.L.: Visual to Sound: Generating Natural Sound for Videos in the Wild (2018). DOI 10.48550/arXiv.1712.01393. URL http://arxiv.org/abs/ 1712.01393. ArXiv:1712.01393 [cs]