Self-supervised object detection from audio-visual correspondence
We tackle the problem of learning object detectors without supervision. Differently from weakly-supervised object detection, we do not assume image-level class labels. Instead, we extract a supervisory signal from audio-visual data, using the audio component to "teach" the object detector. While this problem is related to sound source localisation, it is considerably harder because the detector must classify the objects by type, enumerate each instance of the object, and do so even when the object is silent. We tackle this problem by first designing a self-supervised framework with a contrastive objective that jointly learns to classify and localise objects. Then, without using any supervision, we simply use these self-supervised labels and boxes to train an image-based object detector. With this, we outperform previous unsupervised and weakly-supervised detectors for the task of object detection and sound source localization. We also show that we can align this detector to ground-truth classes with as little as one label per pseudo-class, and show how our method can learn to detect generic objects that go beyond instruments, such as airplanes and cats.
Code (0)
등록된 구현이 없습니다.
Tasks
Objectobject-detectionObject DetectionSound Source LocalizationWeakly Supervised Object DetectionSimilar Papers 제목 키워드 기반
Self-Supervised Learning of Audio-Visual Objects from Video
Our objective is to transform a video into a set of discrete audio-visual objects using self-supervised learning. To this end, we introduce a model that uses attention to localize and group sound sources, and optical flo…
Active Speaker DetectionFace DetectionOptical Flow EstimationSelf-Supervised LearningSelf-supervised Learning of Audio Representations from Audio-Visual Data using Spatial Alignment
Learning from audio-visual data offers many possibilities to express correspondence between the audio and visual content, similar to the human perception that relates aural and visual information. In this work, we presen…
Acoustic Scene ClassificationAction Recognitionobject-detectionObject Detection+4Self-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling
Detecting sound source objects within visual observation is important for autonomous robots to comprehend surrounding environments. Since sounding objects have a large variety with different appearances in our living env…
Self-Supervised LearningSound Source LocalizationSAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
Multimodal deepfakes can exhibit subtle visual artifacts and cross-modal inconsistencies, which remain challenging to detect, especially when detectors are trained primarily on curated synthetic forgeries. Such synthetic…
Self-Supervised LearningDeepFake DetectionThere is More than Meets the Eye: Self-Supervised Multi-Object Detection and Tracking with Sound by Distilling Multimodal Knowledge
Attributes of sound inherent to objects can provide valuable cues to learn rich representations for object detection and tracking. Furthermore, the co-occurrence of audiovisual events in videos can be exploited to locali…
object-detectionObject Detection