paper-with-me

Papers

Class-aware Sounding Objects Localization via Audiovisual Correspondence

2021-12-22 · Di Hu, Yake Wei, Rui Qian, Weiyao Lin, Ruihua Song, Ji-Rong Wen

Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localization without category annotations, i.e., localizing the sounding object and recognizing its category. To address this problem, we propose a two-stage step-by-step learning framework to localize and recognize sounding objects in complex audiovisual scenarios using only the correspondence between audio and vision. First, we propose to determine the sounding area via coarse-grained audiovisual correspondence in the single source cases. Then visual features in the sounding area are leveraged as candidate object representations to establish a category-representation object dictionary for expressive visual character extraction. We generate class-aware object localization maps in cocktail-party scenarios and use audiovisual correspondence to suppress silent areas by referring to this dictionary. Finally, we employ category-level audiovisual consistency as the supervision to achieve fine-grained audio and sounding object distribution alignment. Experiments on both realistic and synthesized videos show that our model is superior in localizing and recognizing objects as well as filtering out silent ones. We also transfer the learned audiovisual network into the unsupervised object detection task, obtaining reasonable performance.

📄 PDF Abstract BibTeX arXiv:2112.11749

Code (1)

gewu-lab/csol_tpami2021 공식 구현 pytorch

Tasks

Objectobject-detectionObject DetectionObject LocalizationUnsupervised Object Detection

Similar Papers 제목 키워드 기반

Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching

2020-10-12 · NeurIPS 2020 12 · Di Hu, Rui Qian, Minyue Jiang, Xiao Tan 외

Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a two-stage learning framework to perform…

ObjectObject Localization

Unveiling Visual Biases in Audio-Visual Localization Benchmarks

2024-08-25 · Liangyu Chen, Zihao Yue, Boshen Xu, Qin Jin

Audio-Visual Source Localization (AVSL) aims to localize the source of sound within a video. In this paper, we identify a significant issue in existing benchmarks: the sounding objects are often easily recognized based s…

audio-visual learningVisual Localization

Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation

2025-09-26 · Jinbae Seo, Hyeongjun Kwon, Kwonyoung Kim, Jiyoung Lee 외 arxiv

Audiovisual instance segmentation (AVIS) requires accurately localizing and tracking sounding objects throughout video sequences. Existing methods suffer from visual bias stemming from two fundamental issues: uniform add…

Instance Segmentation

Cyclic Learning for Binaural Audio Generation and Localization

2024-01-01 · CVPR 2024 1 · Zhaojian Li, Bin Zhao, Yuan Yuan

Binaural audio is obtained by simulating the biological structure of human ears which plays an important role in artificial immersive spaces. A promising approach is to utilize mono audio and corresponding vision to …

Audio GenerationObjectObject Localization

Space-Time Memory Network for Sounding Object Localization in Videos

2021-11-10 · Sizhe Li, Yapeng Tian, Chenliang Xu

Leveraging temporal synchronization and association within sight and sound is an essential step towards robust localization of sounding objects. To this end, we propose a space-time memory network for sounding object loc…

Object Localization