paper-with-me

Papers

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance

2024-12-24 · Yaoyun Zhang, Xuenan Xu, Mengyue Wu

The video-to-audio (V2A) generation task has drawn attention in the field of multimedia due to the practicality in producing Foley sound. Semantic and temporal conditions are fed to the generation model to indicate sound events and temporal occurrence. Recent studies on synthesizing immersive and synchronized audio are faced with challenges on videos with moving visual presence. The temporal condition is not accurate enough while low-resolution semantic condition exacerbates the problem. To tackle these challenges, we propose Smooth-Foley, a V2A generative model taking semantic guidance from the textual label across the generation to enhance both semantic and temporal alignment in audio. Two adapters are trained to leverage pre-trained text-to-audio generation models. A frame adapter integrates high-resolution frame-wise video features while a temporal adapter integrates temporal conditions obtained from similarities of visual frames and textual labels. The incorporation of semantic guidance from textual labels achieves precise audio-video alignment. We conduct extensive quantitative and qualitative experiments. Results show that Smooth-Foley performs better than existing models on both continuous sound scenarios and general scenarios. With semantic guidance, the audio generated by Smooth-Foley exhibits higher quality and better adherence to physical laws.

📄 PDF Abstract BibTeX arXiv:2412.18157

Code (0)

등록된 구현이 없습니다.

Tasks

Audio GenerationVideo Alignment

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Adapter 설명 없음

Similar Papers 제목 키워드 기반

Video-Guided Foley Sound Generation with Multimodal Controls

2024-11-26 · CVPR 2025 1 · Ziyang Chen, Prem Seetharaman, Bryan Russell, Oriol Nieto 외

Generating sound effects for videos often requires creating artistic sound effects that diverge significantly from real-life sources and flexible control in the sound design. To address this problem, we introduce MultiFo…

Audio Generation

Conditional Generation of Audio from Video via Foley Analogies

2023-04-17 · CVPR 2023 1 · Yuexi Du, Ziyang Chen, Justin Salamon, Bryan Russell 외

The sound effects that designers add to videos are designed to convey a particular artistic effect and, thus, may be quite different from a scene's true sound. Inspired by the challenges of creating a soundtrack for a vi…

AutoFoley: Artificial Synthesis of Synchronized Sound Tracks for Silent Videos with Deep Learning

2020-02-21 · Sanchita Ghose, John J. Prevost

In movie productions, the Foley Artist is responsible for creating an overlay soundtrack that helps the movie come alive for the audience. This requires the artist to first identify the sounds that will enhance the exper…

Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound

2024-08-21 · Junwon Lee, Jaekwon Im, Dabin Kim, Juhan Nam

Foley sound synthesis is crucial for multimedia production, enhancing user experience by synchronizing audio and video both temporally and semantically. Recent studies on automating this labor-intensive process through v…

Audio GenerationAudio SynthesisSelf-Supervised LearningVideo/Text-to-Audio Generation+1

FoleyBench: A Benchmark For Video-to-Audio Models

2025-11-17 · Satvik Dixit, Koichi Saito, Zhi Zhong, Yuki Mitsufuji 외 arxiv

Video-to-audio generation (V2A) is of increasing importance in domains such as film post-production, AR/VR, and sound design, particularly for the creation of Foley sound effects synchronized with on-screen actions. Fole…

Audio GenerationVideo Alignment