Multiple Sound Sources Localization from Coarse to Fine
How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this problem, we develop a two-stage audiovisual learning framework that disentangles audio and visual representations of different categories from complex scenes, then performs cross-modal feature alignment in a coarse-to-fine manner. Our model achieves state-of-the-art results on public dataset of localization, as well as considerable performance on multi-source sound localization in complex scenes. We then employ the localization results for sound separation and obtain comparable performance to existing methods. These outcomes demonstrate our model's ability in effectively aligning sounds with specific visual sources. Code is available at https://github.com/shvdiwnkozbw/Multi-Source-Sound-Localization
Code (1)
Similar Papers 제목 키워드 기반
AuralNet: Hierarchical Attention-based 3D Binaural Localization of Overlapping Speakers
We propose AuralNet, a novel 3D multi-source binaural sound source localization approach that localizes overlapping sources in both azimuth and elevation without prior knowledge of the number of sources. AuralNet employs…
Sound Source LocalizationLeveraging Sound Source Trajectories for Universal Sound Separation
Existing methods utilizing spatial information for sound source separation require prior knowledge of the direction of arrival (DOA) of the source or utilize estimated but imprecise localization results, which impairs th…
Sound Source LocalizationDeep Neural Networks for Multiple Speaker Detection and Localization
We propose to use neural networks for simultaneous detection and localization of multiple sound sources in human-robot interaction. In contrast to conventional signal processing techniques, neural network-based sound sou…
Sound Source LocalizationAudio-Visual Grouping Network for Sound Localization from Mixtures
Sound source localization is a typical and challenging task that predicts the location of sound sources in a video. Previous single-source methods mainly used the audio-visual association as clues to localize sounding ob…
Object LocalizationSound Source LocalizationA Proposal-Based Paradigm for Self-Supervised Sound Source Localization in Videos
Humans can easily recognize where and how the sound is produced via watching a scene and listening to corresponding audio cues. To achieve such cross-modal perception on machines, existing methods only use the maps g…
Multiple Instance LearningSound Source Localization