A Proposal-Based Paradigm for Self-Supervised Sound Source Localization in Videos
Humans can easily recognize where and how the sound is produced via watching a scene and listening to corresponding audio cues. To achieve such cross-modal perception on machines, existing methods only use the maps generated by interpolation operations to localize the sound source. As semantic object-level localization is more attractive for potential practical applications, we argue that these existing map-based approaches only provide a coarse-grained and indirect description of the sound source. In this paper, we advocate a novel proposal-based paradigm that can directly perform semantic object-level localization, without any manual annotations. We incorporate the global response map as an unsupervised spatial constraint to weight the proposals according to how well they cover the estimated global shape of the sound source. As a result, our proposal-based sound source localization can be cast into a simpler Multiple Instance Learning (MIL) problem by filtering those instances corresponding to large sound-unrelated regions. Our method achieves state-of-the-art (SOTA) performance when compared to several baselines on multiple datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Multiple Instance LearningSound Source LocalizationSimilar Papers 제목 키워드 기반
Reversing the cycle: self-supervised deep stereo through enhanced monocular distillation
In many fields, self-supervised learning solutions are rapidly evolving and filling the gap with supervised approaches. This fact occurs for depth estimation based on either monocular or stereo, with the latter often pro…
Depth EstimationSelf-Supervised LearningvalidMATPAC++: Enhanced Masked Latent Prediction for Self-Supervised Audio Representation Learning
Masked latent prediction has emerged as a leading paradigm in self-supervised learning (SSL), especially for general audio and music representation learning. While recent methods have demonstrated strong performance, the…
Self-Supervised LearningRepresentation LearningSelf-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling
Detecting sound source objects within visual observation is important for autonomous robots to comprehend surrounding environments. Since sounding objects have a large variety with different appearances in our living env…
Self-Supervised LearningSound Source LocalizationWeakly Supervised Localisation for Fetal Ultrasound Images
This paper addresses the task of detecting and localising fetal anatomical regions in 2D ultrasound images, where only image-level labels are present at training, i.e. without any localisation or segmentation information…
Pose EstimationSegmentationWeakly-supervised Audio-visual Sound Source Detection and Separation
Learning how to localize and separate individual object sounds in the audio channel of the video is a difficult task. Current state-of-the-art methods predict audio masks from artificially mixed spectrograms, known as Mi…
Audio Source SeparationDenoisingObjectSegmentation+3