paper-with-me

Papers

Co-Separating Sounds of Visual Objects

2019-04-16 · ICCV 2019 10 · Ruohan Gao, Kristen Grauman

Learning how objects sound from video is challenging, since they often heavily overlap in a single audio channel. Current methods for visually-guided audio source separation sidestep the issue by training with artificially mixed video clips, but this puts unwieldy restrictions on training data collection and may even prevent learning the properties of "true" mixed sounds. We introduce a co-separation training paradigm that permits learning object-level sounds from unlabeled multi-source videos. Our novel training objective requires that the deep neural network's separated audio for similar-looking objects be consistently identifiable, while simultaneously reproducing accurate video-level audio tracks for each source training pair. Our approach disentangles sounds in realistic test videos, even in cases where an object was not observed individually during training. We obtain state-of-the-art results on visually-guided audio source separation and audio denoising for the MUSIC, AudioSet, and AV-Bench datasets.

📄 PDF Abstract BibTeX arXiv:1904.07750

Code (3)

YashNita/Co-Separating-Sound-Object- pytorch
manhnguyen1998/co_separation_encoder_decoder pytorch
rhgao/co-separation pytorch

Tasks

Audio DenoisingAudio Source SeparationDenoising

Similar Papers 제목 키워드 기반

Learning to Separate Object Sounds by Watching Unlabeled Video

2018-04-05 · ECCV 2018 9 · Ruohan Gao, Rogerio Feris, Kristen Grauman

Perceiving a scene most fully requires all the senses. Yet modeling how objects look and sound is challenging: most natural scenes and events contain multiple objects, and the audio track mixes all the sound sources toge…

Audio DenoisingAudio Source SeparationDenoisingMulti-Label Learning

The Sound of Motions

2019-04-11 · ICCV 2019 10 · Hang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio Torralba

Sounds originate from object motions and vibrations of surrounding air. Inspired by the fact that humans is capable of interpreting sound sources from how objects move visually, we propose a novel system that explicitly …

Separating Invisible Sounds Toward Universal Audiovisual Scene-Aware Sound Separation

2023-10-18 · Yiyang Su, Ali Vosoughi, Shijian Deng, Yapeng Tian 외

The audio-visual sound separation field assumes visible sources in videos, but this excludes invisible sounds beyond the camera's view. Current methods struggle with such sounds lacking visible cues. This paper introduce…

cross-modal alignment

Self-Supervised Audio-Visual Co-Segmentation

2019-04-18 · Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh Mcdermott 외

Segmenting objects in images and separating sound sources in audio are challenging tasks, in part because traditional approaches require large amounts of labeled data. In this paper we develop a neural network model for …

Image SegmentationSegmentationSemantic Segmentation

RealImpact: A Dataset of Impact Sound Fields for Real Objects

2023-06-16 · CVPR 2023 1 · Samuel Clarke, Ruohan Gao, Mason Wang, Mark Rau 외

Objects make unique sounds under different perturbations, environment conditions, and poses relative to the listener. While prior works have modeled impact sounds and sound propagation in simulation, we lack a standard d…

audio-visual learning