paper-with-me

Papers

Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes

2024-07-15 · Yaoting Wang, Peiwen Sun, Dongzhan Zhou, Guangyao Li, Honggang Zhang, Di Hu

Traditional reference segmentation tasks have predominantly focused on silent visual scenes, neglecting the integral role of multimodal perception and interaction in human experiences. In this work, we introduce a novel task called Reference Audio-Visual Segmentation (Ref-AVS), which seeks to segment objects within the visual domain based on expressions containing multimodal cues. Such expressions are articulated in natural language forms but are enriched with multimodal cues, including audio and visual descriptions. To facilitate this research, we construct the first Ref-AVS benchmark, which provides pixel-level annotations for objects described in corresponding multimodal-cue expressions. To tackle the Ref-AVS task, we propose a new method that adequately utilizes multimodal cues to offer precise segmentation guidance. Finally, we conduct quantitative and qualitative experiments on three test subsets to compare our approach with existing methods from related tasks. The results demonstrate the effectiveness of our method, highlighting its capability to precisely segment objects using multimodal-cue expressions. Dataset is available at \href{https://gewu-lab.github.io/Ref-AVS}{https://gewu-lab.github.io/Ref-AVS}.

📄 PDF Abstract BibTeX arXiv:2407.10957

Code (0)

등록된 구현이 없습니다.

Tasks

Segmentation

Similar Papers 제목 키워드 기반

TSAM: Temporal SAM Augmented with Multimodal Prompts for Referring Audio-Visual Segmentation

2025-01-01 · CVPR 2025 1 · Abduljalil Radman, Jorma Laaksonen

Referring audio-visual segmentation (Ref-AVS) aims to segment objects within audio-visual scenes using multimodal cues embedded in text expressions. While the Segment Anything Model (SAM) has revolutionized visual se…

Referring Audio-Visual Segmentation

Can Textual Semantics Mitigate Sounding Object Segmentation Preference?

2024-07-15 · Yaoting Wang, Peiwen Sun, Yuanchao Li, Honggang Zhang 외

The Audio-Visual Segmentation (AVS) task aims to segment sounding objects in the visual space using audio cues. However, in this work, it is recognized that previous AVS methods show a heavy reliance on detrimental segme…

Language ModellingLarge Language ModelObjectSegmentation+1

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes

2025-06-02 · CVPR 2025 1 · Yuji Wang, Haoran Xu, Yong liu, Jiaze Li 외

Reference Audio-Visual Segmentation (Ref-AVS) aims to provide a pixel-wise scene understanding in Language-aided Audio-Visual Scenes (LAVS). This task requires the model to continuously segment objects referred to by tex…

Scene Understanding

Multimodal Referring Segmentation: A Survey

2025-08-01 · Henghui Ding, Song Tang, Shuting He, Chang Liu 외 arxiv

Multimodal referring segmentation aims to segment target objects in visual scenes, such as images, videos, and 3D scenes, based on referring expressions in text or audio format. This task plays a crucial role in practica…

Referring Expression

3D Audio-Visual Segmentation

2024-11-04 · Artem Sokolov, Swapnil Bhosale, Xiatian Zhu

Recognizing the sounding objects in scenes is a longstanding objective in embodied AI, with diverse applications in robotics and AR/VR/MR. To that end, Audio-Visual Segmentation (AVS), taking as condition an audio signal…

Segmentation