paper-with-me

Papers

Separating Invisible Sounds Toward Universal Audiovisual Scene-Aware Sound Separation

2023-10-18 · Yiyang Su, Ali Vosoughi, Shijian Deng, Yapeng Tian, Chenliang Xu

The audio-visual sound separation field assumes visible sources in videos, but this excludes invisible sounds beyond the camera's view. Current methods struggle with such sounds lacking visible cues. This paper introduces a novel "Audio-Visual Scene-Aware Separation" (AVSA-Sep) framework. It includes a semantic parser for visible and invisible sounds and a separator for scene-informed separation. AVSA-Sep successfully separates both sound types, with joint training and cross-modal alignment enhancing effectiveness.

📄 PDF Abstract BibTeX arXiv:2310.11713

Code (0)

등록된 구현이 없습니다.

Tasks

cross-modal alignment

Similar Papers 제목 키워드 기반

Large Scale Audiovisual Learning of Sounds with Weakly Labeled Data

2020-05-29 · Haytham M. Fayek, Anurag Kumar

Recognizing sounds is a key aspect of computational audio scene analysis and machine perception. In this paper, we advocate that sound recognition is inherently a multi-modal audiovisual task in that it is easier to diff…

Audio Classification

Audio Prompt Tuning for Universal Sound Separation

2023-11-30 · Yuzhuo Liu, Xubo Liu, Yan Zhao, Yuanyuan Wang 외

Universal sound separation (USS) is a task to separate arbitrary sounds from an audio mixture. Existing USS systems are capable of separating arbitrary sources, given a few examples of the target sources as queries. Howe…

Cross-Task Transfer for Geotagged Audiovisual Aerial Scene Recognition

2020-05-18 · ECCV 2020 8 · Di Hu, Xuhong LI, Lichao Mou, Pu Jin 외

Aerial scene recognition is a fundamental task in remote sensing and has recently received increased interest. While the visual information from overhead images with powerful models and efficient algorithms yields consid…

Scene Recognition

Universal Sound Separation

2019-05-08 · Ilya Kavalerov, Scott Wisdom, Hakan Erdogan, Brian Patton 외

Recent deep learning approaches have achieved impressive performance on speech enhancement and separation tasks. However, these approaches have not been investigated for separating mixtures of arbitrary sounds of differe…

Speech EnhancementSpeech Separation

Learning to Separate Object Sounds by Watching Unlabeled Video

2018-04-05 · ECCV 2018 9 · Ruohan Gao, Rogerio Feris, Kristen Grauman

Perceiving a scene most fully requires all the senses. Yet modeling how objects look and sound is challenging: most natural scenes and events contain multiple objects, and the audio track mixes all the sound sources toge…

Audio DenoisingAudio Source SeparationDenoisingMulti-Label Learning