paper-with-me

홈 › Papers

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

2026-08-04 · Sandy Abdo, Bill Kapralos, Priyamvada Tripathi, KC Collins, Adam Dubrowski arxiv

Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability. Artificial intelligence (AI)-driven audio generative models are rapidly growing in popularity and have the potential to transform the way sound is synthesized and used across various applications. In response to this growing momentum, this chapter reviews and analyzes recent AI-based generative models for sound effect synthesis, with a focus on how different input modalities (text, visual, audio, and multimodal) affect the quality, controllability, and contextual relevance of the generated audio. It examines 30 peer-reviewed articles sourced from Google Scholar, IEEE Xplore, and the ACM Digital Library, exploring the evolution of AI generative models over the past five years. The results show that multiple models achieved state-of-the-art performance, producing high-fidelity, semantically aligned, and increasingly temporally coherent sound effects across tasks. However, despite these advances, the review identifies persistent challenges, including limitations in temporal synchronization for complex multi-event scenarios, gaps between objective metrics and human perception, and trade-offs between controllability and generative diversity. Overall, the chapter highlights that AI-driven sound effect generation is progressing toward more adaptive, scalable, and context-aware systems, offering significant implications for future sound design workflows and interactive media applications.

📄 PDF Abstract BibTeX arXiv:2608.03742

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Retrieval: Generating Narratives in Conversational Recommender Systems

2024-10-22 · Krishna Sayana, Raghavendra Vasudeva, Yuri Vasilevski, Kun Su 외

The recent advances in Large Language Model's generation and reasoning capabilities present an opportunity to develop truly conversational recommendation systems. However, effectively integrating recommender system knowl…

Conversational RecommendationRecommendation SystemsRetrievalText Generation

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions

2026-06-01 · Jiashuo Yu, Yao Yao, Boyu Chen, Alex Wang arxiv

We address the challenge of generating high-fidelity, long-form soundtracks that remain coherent across scene transitions. Existing AI music systems are mainly designed for short, isolated clips and lack mechanisms to en…

Audio Generation

Human-in-the-Loop for Data Collection: a Multi-Target Counter Narrative Dataset to Fight Online Hate Speech

2021-07-19 · ACL 2021 5 · Margherita Fanton, Helena Bonaldi, Serra Sinem Tekiroglu, Marco Guerini

Undermining the impact of hateful content with informed and non-aggressive responses, called counter narratives, has emerged as a possible solution for having healthier online communities. Thus, some NLP studies have sta…

Language ModelingLanguage Modelling

REGEN: A Dataset and Benchmarks with Natural Language Critiques and Narratives

2025-03-14 · Kun Su, Krishna Sayana, Hubert Pham, James Pine 외

This paper introduces a novel dataset REGEN (Reviews Enhanced with GEnerative Narratives), designed to benchmark the conversational capabilities of recommender Large Language Models (LLMs), addressing the limitations of …

Conversational Recommendation

AudioStory: Generating Long-Form Narrative Audio with Large Language Models

2025-08-27 · Yuxin Guo, Teng Wang, Yuying Ge, Shijie Ma 외 arxiv

Recent advances in text-to-audio (TTA) generation excel at synthesizing short audio clips but struggle with long-form narrative audio, which requires temporal coherence and compositional reasoning. To address this gap, w…

Audio Generation