Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching
Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a two-stage learning framework to perform self-supervised class-aware sounding object localization. First, we propose to learn robust object representations by aggregating the candidate sound localization results in the single source scenes. Then, class-aware object localization maps are generated in the cocktail-party scenarios by referring the pre-learned object knowledge, and the sounding objects are accordingly selected by matching audio and visual object category distributions, where the audiovisual consistency is viewed as the self-supervised signal. Experimental results in both realistic and synthesized cocktail-party videos demonstrate that our model is superior in filtering out silent objects and pointing out the location of sounding objects of different classes. Code is available at https://github.com/DTaoo/Discriminative-Sounding-Objects-Localization.
Code (1)
Tasks
ObjectObject LocalizationSimilar Papers 제목 키워드 기반
Class-aware Sounding Objects Localization via Audiovisual Correspondence
Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localiza…
Objectobject-detectionObject DetectionObject Localization+1Multi-scale Multi-instance Visual Sound Localization and Segmentation
Visual sound localization is a typical and challenging problem that predicts the location of objects corresponding to the sound source in a video. Previous methods mainly used the audio-visual association between global …
Object LocalizationSelf-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling
Detecting sound source objects within visual observation is important for autonomous robots to comprehend surrounding environments. Since sounding objects have a large variety with different appearances in our living env…
Self-Supervised LearningSound Source LocalizationHearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization
Large-scale vision-language models demonstrate strong multimodal alignment and generalization across diverse tasks. Among them, CLIP stands out as one of the most successful approaches. In this work, we extend the applic…
Scene UnderstandingSound Source LocalizationSpace-Time Memory Network for Sounding Object Localization in Videos
Leveraging temporal synchronization and association within sight and sound is an essential step towards robust localization of sounding objects. To this end, we propose a space-time memory network for sounding object loc…
Object Localization