paper-with-me

Papers

Self-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling

2020-07-28 · Yoshiki Masuyama, Yoshiaki Bando, Kohei Yatabe, Yoko Sasaki, Masaki Onishi, Yasuhiro Oikawa

Detecting sound source objects within visual observation is important for autonomous robots to comprehend surrounding environments. Since sounding objects have a large variety with different appearances in our living environments, labeling all sounding objects is impossible in practice. This calls for self-supervised learning which does not require manual labeling. Most of conventional self-supervised learning uses monaural audio signals and images and cannot distinguish sound source objects having similar appearances due to poor spatial information in audio signals. To solve this problem, this paper presents a self-supervised training method using 360{\deg} images and multichannel audio signals. By incorporating with the spatial information in multichannel audio signals, our method trains deep neural networks (DNNs) to distinguish multiple sound source objects. Our system for localizing sound source objects in the image is composed of audio and visual DNNs. The visual DNN is trained to localize sound source candidates within an input image. The audio DNN verifies whether each candidate actually produces sound or not. These DNNs are jointly trained in a self-supervised manner based on a probabilistic spatial audio model. Experimental results with simulated data showed that the DNNs trained by our method localized multiple speakers. We also demonstrate that the visual DNN detected objects including talking visitors and specific exhibits from real data recorded in a science museum.

📄 PDF Abstract BibTeX arXiv:2007.13976

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningSound Source Localization

Similar Papers 제목 키워드 기반

Learning from Silence and Noise for Visual Sound Source Localization

2025-08-29 · Xavier Juanola, Giovana Morais, Magdalena Fuentes, Gloria Haro arxiv

Visual sound source localization is a fundamental perception task that aims to detect the location of sounding sources in a video given its audio. Despite recent progress, we identify two shortcomings in current methods:…

Sound Source LocalizationSemantic correspondenceCross-Modal Retrieval

Visually Guided Sound Source Separation and Localization using Self-Supervised Motion Representations

2021-04-17 · Lingyu Zhu, Esa Rahtu

The objective of this paper is to perform audio-visual sound source separation, i.e.~to separate component audios from a mixture based on the videos of sound sources. Moreover, we aim to pinpoint the source location in t…

Optical Flow EstimationVisually Guided Sound Source Separation

Self-Supervised Audio-Visual Co-Segmentation

2019-04-18 · Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh Mcdermott 외

Segmenting objects in images and separating sound sources in audio are challenging tasks, in part because traditional approaches require large amounts of labeled data. In this paper we develop a neural network model for …

Image SegmentationSegmentationSemantic Segmentation

Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source Localization

2023-08-09 · Tianyu Liu, Peng Zhang, Wei Huang, Yufei zha 외

Self-supervised sound source localization is usually challenged by the modality inconsistency. In recent studies, contrastive learning based strategies have shown promising to establish such a consistent correspondence b…

Contrastive LearningSound Source Localization

Telling Left from Right: Learning Spatial Correspondence of Sight and Sound

2020-06-11 · CVPR 2020 6 · Karren Yang, Bryan Russell, Justin Salamon

Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic informa…

audio-visual learning