paper-with-me

Papers

Audio-Visual Segmentation

2022-07-11 · Jinxing Zhou, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, Yiran Zhong

We propose to explore a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame. To facilitate this research, we construct the first audio-visual segmentation benchmark (AVSBench), providing pixel-wise annotations for the sounding objects in audible videos. Two settings are studied with this benchmark: 1) semi-supervised audio-visual segmentation with a single sound source and 2) fully-supervised audio-visual segmentation with multiple sound sources. To deal with the AVS problem, we propose a novel method that uses a temporal pixel-wise audio-visual interaction module to inject audio semantics as guidance for the visual segmentation process. We also design a regularization loss to encourage the audio-visual mapping during training. Quantitative and qualitative experiments on the AVSBench compare our approach to several existing methods from related tasks, demonstrating that the proposed method is promising for building a bridge between the audio and pixel-wise visual semantics. Code is available at https://github.com/OpenNLPLab/AVSBench.

📄 PDF Abstract BibTeX arXiv:2207.05042

Code (2)

opennlplab/avsbench 공식 구현 pytorch
Sunjuhyeong/SAM_STBAVA pytorch

Tasks

Segmentation

Similar Papers 제목 키워드 기반

Weakly-Supervised Audio-Visual Segmentation

2023-11-25 · NeurIPS 2023 11

Audio-visual segmentation is a challenging task that aims to predict pixel-level masks for sound sources in a video. Previous work applied a comprehensive manually designed architecture with countless pixel-wise accurate…

Contrastive LearningSegmentation

AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation

2023-05-03 · Shentong Mo, Yapeng Tian

Segment Anything Model (SAM) has recently shown its powerful effectiveness in visual segmentation tasks. However, there is less exploration concerning how SAM works on audio-visual tasks, such as visual sound localizatio…

DecoderObject LocalizationSegmentationVisual Localization

Audio-Visual Segmentation with Semantics

2023-01-30 · Jinxing Zhou, Xuyang Shen, Jianyuan Wang, Jiayi Zhang 외

We propose a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame. To facilitate this research, we constr…

SegmentationSemantic SegmentationVideo Semantic Segmentation

BAVS: Bootstrapping Audio-Visual Segmentation by Integrating Foundation Knowledge

2023-08-20 · Chen Liu, Peike Li, Hu Zhang, Lincheng Li 외

Given an audio-visual pair, audio-visual segmentation (AVS) aims to locate sounding sources by predicting pixel-wise maps. Previous methods assume that each sound component in an audio signal always has a visual counterp…

Audio ClassificationSegmentation

Unraveling Instance Associations: A Closer Look for Audio-Visual Segmentation

2023-04-06 · CVPR 2024 1 · Yuanhong Chen, Yuyuan Liu, Hu Wang, Fengbei Liu 외

Audio-visual segmentation (AVS) is a challenging task that involves accurately segmenting sounding objects based on audio-visual cues. The effectiveness of audio-visual learning critically depends on achieving accurate c…

audio-visual learningContrastive Learningcross-modal alignmentSegmentation