paper-with-me

Papers

Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos

2024-06-13 · Changan Chen, Puyuan Peng, Ami Baid, Zihui Xue, Wei-Ning Hsu, David Harwath, Kristen Grauman

Generating realistic audio for human actions is important for many applications, such as creating sound effects for films or virtual reality games. Existing approaches implicitly assume total correspondence between the video and audio during training, yet many sounds happen off-screen and have weak to no correspondence with the visuals -- resulting in uncontrolled ambient sounds or hallucinations at test time. We propose a novel ambient-aware audio generation model, AV-LDM. We devise a novel audio-conditioning mechanism to learn to disentangle foreground action sounds from the ambient background sounds in in-the-wild training videos. Given a novel silent video, our model uses retrieval-augmented generation to create audio that matches the visual content both semantically and temporally. We train and evaluate our model on two in-the-wild egocentric video datasets, Ego4D and EPIC-KITCHENS, and we introduce Ego4D-Sounds -- 1.2M curated clips with action-audio correspondence. Our model outperforms an array of existing methods, allows controllable generation of the ambient sound, and even shows promise for generalizing to computer graphics game clips. Overall, our approach is the first to focus video-to-audio generation faithfully on the observed visual content despite training from uncurated clips with natural background sounds.

📄 PDF Abstract BibTeX arXiv:2406.09272

Code (0)

등록된 구현이 없습니다.

Tasks

Audio GenerationRetrieval-augmented Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MetaBGM: Dynamic Soundtrack Transformation For Continuous Multi-Scene Experiences With Ambient Awareness And Personalization

2024-09-05 · Haoxuan Liu, ZiHao Wang, HaoRong Hong, Youwei Feng 외

This paper introduces MetaBGM, a groundbreaking framework for generating background music that adapts to dynamic scenes and real-time user interactions. We define multi-scene as variations in environmental contexts, such…

Audio Generation

Geometrically-Motivated Primary-Ambient Decomposition With Center-Channel Extraction

2022-06-05 · Jouni Paulus, Matteo Torcoli

A geometrically-motivated method for primary-ambient decomposition is proposed and evaluated in an up-mixing application. The method consists of two steps, accommodating a particularly intuitive explanation. The first st…

MR4MR: Mixed Reality for Melody Reincarnation

2022-09-15 · Atsuya Kobayashi, Ryogo Ishino, Ryuku Nobusue, Takumi Inoue 외

There is a long history of an effort made to explore musical elements with the entities and spaces around us, such as musique concr\`ete and ambient music. In the context of computer music and digital art, interactive ex…

Mixed RealityMusic Generation

Read the Room: Adapting a Robot's Voice to Ambient and Social Contexts

2022-05-10 · Paige Tuttosi, Emma Hughson, Akihiro Matsufuji, Angelica Lim

How should a robot speak in a formal, quiet and dark, or a bright, lively and noisy environment? By designing robots to speak in a more social and ambient-appropriate manner we can improve perceived awareness and intelli…

Speech SynthesisVoice Conversion

Foreground-Background Ambient Sound Scene Separation

2020-05-11 · Michel Olvera, Emmanuel Vincent, Romain Serizel, Gilles Gasso

Ambient sound scenes typically comprise multiple short events occurring on top of a somewhat stationary background. We consider the task of separating these events from the background, which we call foreground-background…