paper-with-me

Papers

Dual Normalization Multitasking for Audio-Visual Sounding Object Localization

2021-06-01 · Tokuhiro Nishikawa, Daiki Shimada, Jerry Jun Yokono

Although several research works have been reported on audio-visual sound source localization in unconstrained videos, no datasets and metrics have been proposed in the literature to quantitatively evaluate its performance. Defining the ground truth for sound source localization is difficult, because the location where the sound is produced is not limited to the range of the source object, but the vibrations propagate and spread through the surrounding objects. Therefore we propose a new concept, Sounding Object, to reduce the ambiguity of the visual location of sound, making it possible to annotate the location of the wide range of sound sources. With newly proposed metrics for quantitative evaluation, we formulate the problem of Audio-Visual Sounding Object Localization (AVSOL). We also created the evaluation dataset (AVSOL-E dataset) by manually annotating the test set of well-known Audio-Visual Event (AVE) dataset. To tackle this new AVSOL problem, we propose a novel multitask training strategy and architecture called Dual Normalization Multitasking (DNM), which aggregates the Audio-Visual Correspondence (AVC) task and the classification task for video events into a single audio-visual similarity map. By efficiently utilize both supervisions by DNM, our proposed architecture significantly outperforms the baseline methods.

📄 PDF Abstract BibTeX arXiv:2106.00180

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectObject LocalizationSound Source Localization

Similar Papers 제목 키워드 기반

Cyclic Co-Learning of Sounding Object Visual Grounding and Sound Separation

2021-04-05 · CVPR 2021 1 · Yapeng Tian, Di Hu, Chenliang Xu

There are rich synchronized audio and visual events in our daily life. Inside the events, audio scenes are associated with the corresponding visual objects; meanwhile, sounding objects can indicate and help to separate t…

ObjectVisual Grounding

Class-aware Sounding Objects Localization via Audiovisual Correspondence

2021-12-22 · Di Hu, Yake Wei, Rui Qian, Weiyao Lin 외

Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localiza…

Objectobject-detectionObject DetectionObject Localization+1

Audio-Visual Segmentation by Exploring Cross-Modal Mutual Semantics

2023-07-31 · Chen Liu, Peike Li, Xingqun Qi, Hu Zhang 외

The audio-visual segmentation (AVS) task aims to segment sounding objects from a given video. Existing works mainly focus on fusing audio and visual features of a given video to achieve sounding object masks. However, we…

ObjectSegmentationSemantic Segmentation

Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation

2023-09-18 · Shaofei Huang, Han Li, Yuqing Wang, Hongji Zhu 외

Audio visual segmentation (AVS) aims to segment the sounding objects for each frame of a given video. To distinguish the sounding objects from silent ones, both audio-visual semantic correspondence and temporal interacti…

ObjectSemantic correspondence

Can Textual Semantics Mitigate Sounding Object Segmentation Preference?

2024-07-15 · Yaoting Wang, Peiwen Sun, Yuanchao Li, Honggang Zhang 외

The Audio-Visual Segmentation (AVS) task aims to segment sounding objects in the visual space using audio cues. However, in this work, it is recognized that previous AVS methods show a heavy reliance on detrimental segme…

Language ModellingLarge Language ModelObjectSegmentation+1