paper-with-me

Papers

Weakly-supervised Audio-visual Sound Source Detection and Separation

2021-03-25 · Tanzila Rahman, Leonid Sigal

Learning how to localize and separate individual object sounds in the audio channel of the video is a difficult task. Current state-of-the-art methods predict audio masks from artificially mixed spectrograms, known as Mix-and-Separate framework. We propose an audio-visual co-segmentation, where the network learns both what individual objects look and sound like, from videos labeled with only object labels. Unlike other recent visually-guided audio source separation frameworks, our architecture can be learned in an end-to-end manner and requires no additional supervision or bounding box proposals. Specifically, we introduce weakly-supervised object segmentation in the context of sound separation. We also formulate spectrogram mask prediction using a set of learned mask bases, which combine using coefficients conditioned on the output of object segmentation , a design that facilitates separation. Extensive experiments on the MUSIC dataset show that our proposed approach outperforms state-of-the-art methods on visually guided sound source separation and sound denoising.

📄 PDF Abstract BibTeX arXiv:2104.02606

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Source SeparationDenoisingObjectSegmentationSemantic SegmentationVisually Guided Sound Source SeparationWeakly-Supervised Object Segmentation

Similar Papers 제목 키워드 기반

Weakly-Supervised Audio-Visual Segmentation

2023-11-25 · NeurIPS 2023 11

Audio-visual segmentation is a challenging task that aims to predict pixel-level masks for sound sources in a video. Previous work applied a comprehensive manually designed architecture with countless pixel-wise accurate…

Contrastive LearningSegmentation

Localize to Binauralize: Audio Spatialization From Visual Sound Source Localization

2021-01-01 · ICCV 2021 10 · Kranthi Kumar Rachavarapu, Aakanksha, Vignesh Sundaresha, A. N. Rajagopalan

Videos with binaural audios provide an immersive viewing experience by enabling 3D sound sensation. Recent works attempt to generate binaural audio in a multimodal learning framework using large quantities of videos …

Audio GenerationSound Source Localization

Can audio-visual integration strengthen robustness under multimodal attacks?

2021-04-05 · CVPR 2021 1 · Yapeng Tian, Chenliang Xu

In this paper, we propose to make a systematic study on machines multisensory perception under attacks. We use the audio-visual event recognition task against multimodal adversarial attacks as a proxy to investigate the …

audio-visual learningVisual Localization

T-VSL: Text-Guided Visual Sound Source Localization in Mixtures

2024-04-02 · CVPR 2024 1 · Tanvir Mahmud, Yapeng Tian, Diana Marculescu

Visual sound source localization poses a significant challenge in identifying the semantic region of each sounding source within a video. Existing self-supervised and weakly supervised source localization methods struggl…

Sound Source Localization

A Closer Look at Weakly-Supervised Audio-Visual Source Localization

2022-08-30 · Shentong Mo, Pedro Morgado

Audio-visual source localization is a challenging task that aims to predict the location of visual sound sources in a video. Since collecting ground-truth annotations of sounding objects can be costly, a plethora of weak…

Sound Source Localization