paper-with-me

Papers

AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer

2026-03-16 · Pengjun Fang, Yingqing He, Yazhou Xing, Qifeng Chen, Ser-Nam Lim, Harry Yang arxiv

Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data, such as conflating acoustically distinct sounds under coarse labels, and textual ambiguity in describing micro-acoustic features. These bottlenecks make it difficult to perform fine-grained sound synthesis using text-controlled modes. To address these limitations, we propose AC-Foley, an audio-conditioned V2A model that directly leverages reference audio to achieve precise and fine-grained control over generated sounds. This approach enables fine-grained sound synthesis, timbre transfer, zero-shot sound generation, and improved audio quality. By directly conditioning on audio signals, our approach bypasses the semantic ambiguities of text descriptions while enabling precise manipulation of acoustic attributes. Empirically, AC-Foley achieves state-of-the-art performance for Foley generation when conditioned on reference audio, while remaining competitive with state-of-the-art video-to-audio methods even without audio conditioning. Code and demo are available at: https://ff2416.github.io/AC-Foley-Page

📄 PDF Abstract BibTeX arXiv:2603.15597

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Video-Guided Foley Sound Generation with Multimodal Controls

2024-11-26 · CVPR 2025 1 · Ziyang Chen, Prem Seetharaman, Bryan Russell, Oriol Nieto 외

Generating sound effects for videos often requires creating artistic sound effects that diverge significantly from real-life sources and flexible control in the sound design. To address this problem, we introduce MultiFo…

Audio Generation

Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video

2025-10-24 · Ciara Rowles, Varun Jampani, Simon Donné, Shimon Vainer 외 arxiv

Foley Control is a lightweight approach to video-guided Foley that keeps pretrained single-modality models frozen and learns only a small cross-attention bridge between them. We connect V-JEPA2 video embeddings to a froz…

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling

2026-04-16 · Jianxuan Yang, Xinyue Guo, Zhi Cheng, Kai Wang 외 arxiv

Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-grained controllability remains challenging. Existing methods suffer from weak tex…

Audio Generation

FoleyGenEx: Unified Video-to-Audio Generation with Multi-Modal Control, Temporal Alignment, and Semantic Precision

2026-06-12 · Shiyao Wang, Xijuan Zeng, Hui Wang, Shiwan Zhao 외 arxiv

We present FoleyGenEx, a unified video-to-audio (VTA) framework integrating multi-modal control, frame-level temporal alignment, and fine-grained semantics, enabling synchronized, versatile audio synthesis for diverse ta…

Data AugmentationAudio Generation

FoleyBench: A Benchmark For Video-to-Audio Models

2025-11-17 · Satvik Dixit, Koichi Saito, Zhi Zhong, Yuki Mitsufuji 외 arxiv

Video-to-audio generation (V2A) is of increasing importance in domains such as film post-production, AR/VR, and sound design, particularly for the creation of Foley sound effects synchronized with on-screen actions. Fole…

Audio GenerationVideo Alignment