paper-with-me

Papers

Audio-Visual Segmentation by Exploring Cross-Modal Mutual Semantics

2023-07-31 · Chen Liu, Peike Li, Xingqun Qi, Hu Zhang, Lincheng Li, Dadong Wang, Xin Yu

The audio-visual segmentation (AVS) task aims to segment sounding objects from a given video. Existing works mainly focus on fusing audio and visual features of a given video to achieve sounding object masks. However, we observed that prior arts are prone to segment a certain salient object in a video regardless of the audio information. This is because sounding objects are often the most salient ones in the AVS dataset. Thus, current AVS methods might fail to localize genuine sounding objects due to the dataset bias. In this work, we present an audio-visual instance-aware segmentation approach to overcome the dataset bias. In a nutshell, our method first localizes potential sounding objects in a video by an object segmentation network, and then associates the sounding object candidates with the given audio. We notice that an object could be a sounding object in one video but a silent one in another video. This would bring ambiguity in training our object segmentation network as only sounding objects have corresponding segmentation masks. We thus propose a silent object-aware segmentation objective to alleviate the ambiguity. Moreover, since the category information of audio is unknown, especially for multiple sounding sources, we propose to explore the audio-visual semantic correlation and then associate audio with potential objects. Specifically, we attend predicted audio category scores to potential instance masks and these scores will highlight corresponding sounding instances while suppressing inaudible ones. When we enforce the attended instance masks to resemble the ground-truth mask, we are able to establish audio-visual semantics correlation. Experimental results on the AVS benchmarks demonstrate that our method can effectively segment sounding objects without being biased to salient objects.

📄 PDF Abstract BibTeX arXiv:2307.16620

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectSegmentationSemantic Segmentation

Methods 이 논문이 사용한 방법론

fail 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation

2025-01-14 · Sitong Gong, Yunzhi Zhuge, Lu Zhang, Yifan Wang 외

The essence of audio-visual segmentation (AVS) lies in locating and delineating sound-emitting objects within a video stream. While Transformer-based methods have shown promise, their handling of long-range dependencies …

MambaVideo Understanding

Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video Parsing

2021-12-01 · NeurIPS 2021 12 · Yan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin 외

The audio-visual video parsing task aims to temporally parse a video into audio or visual event categories. However, it is labor intensive to temporally annotate audio and visual events and thus hampers the learning of a…

Expression Prompt Collaboration Transformer for Universal Referring Video Object Segmentation

2023-08-08 · Jiajun Chen, Jiacheng Lin, Guojin Zhong, Haolong Fu 외

Audio-guided Video Object Segmentation (A-VOS) and Referring Video Object Segmentation (R-VOS) are two highly related tasks that both aim to segment specific objects from video sequences according to expression prompts. …

Contrastive LearningObjectReferring Expression SegmentationReferring Video Object Segmentation+4

AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation

2023-05-03 · Shentong Mo, Yapeng Tian

Segment Anything Model (SAM) has recently shown its powerful effectiveness in visual segmentation tasks. However, there is less exploration concerning how SAM works on audio-visual tasks, such as visual sound localizatio…

DecoderObject LocalizationSegmentationVisual Localization

Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation

2023-05-25 · Shilin Yan, Renrui Zhang, Ziyu Guo, Wenchao Chen 외

Recently, video object segmentation (VOS) referred by multi-modal signals, e.g., language and audio, has evoked increasing attention in both industry and academia. It is challenging for exploring the semantic alignment w…

ObjectReferring Expression SegmentationReferring Video Object SegmentationSemantic Segmentation+2