paper-with-me

Papers

Spectrum-guided Multi-granularity Referring Video Object Segmentation

2023-07-25 · ICCV 2023 1 · Bo Miao, Mohammed Bennamoun, Yongsheng Gao, Ajmal Mian

Current referring video object segmentation (R-VOS) techniques extract conditional kernels from encoded (low-resolution) vision-language features to segment the decoded high-resolution features. We discovered that this causes significant feature drift, which the segmentation kernels struggle to perceive during the forward computation. This negatively affects the ability of segmentation kernels. To address the drift problem, we propose a Spectrum-guided Multi-granularity (SgMg) approach, which performs direct segmentation on the encoded features and employs visual details to further optimize the masks. In addition, we propose Spectrum-guided Cross-modal Fusion (SCF) to perform intra-frame global interactions in the spectral domain for effective multimodal representation. Finally, we extend SgMg to perform multi-object R-VOS, a new paradigm that enables simultaneous segmentation of multiple referred objects in a video. This not only makes R-VOS faster, but also more practical. Extensive experiments show that SgMg achieves state-of-the-art performance on four video benchmark datasets, outperforming the nearest competitor by 2.8% points on Ref-YouTube-VOS. Our extended SgMg enables multi-object R-VOS, runs about 3 times faster while maintaining satisfactory performance. Code is available at https://github.com/bo-miao/SgMg.

📄 PDF Abstract BibTeX arXiv:2307.13537

Code (1)

bo-miao/sgmg 공식 구현 pytorch

Tasks

ObjectReferring Expression SegmentationReferring Video Object SegmentationSegmentationSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation

Similar Papers 제목 키워드 기반

Multi-Level Representation Learning With Semantic Alignment for Referring Video Object Segmentation

2022-01-01 · CVPR 2022 1 · Dongming Wu, Xingping Dong, Ling Shao, Jianbing Shen

Referring video object segmentation (RVOS) is a challenging language-guided video grounding task, which requires comprehensively understanding the semantic information of both video content and language queries for o…

ObjectReferring Expression SegmentationReferring Video Object SegmentationRepresentation Learning+5

Deeply Interleaved Two-Stream Encoder for Referring Video Segmentation

2022-03-30 · Guang Feng, Lihe Zhang, Zhiwei Hu, Huchuan Lu

Referring video segmentation aims to segment the corresponding video object described by the language expression. To address this task, we first design a two-stream encoder to extract CNN-based visual features and transf…

Referring Expression SegmentationVideo SegmentationVideo Semantic SegmentationVocal Bursts Valence Prediction

MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation

2025-12-11 · Henghui Ding, Chang Liu, Shuting He, Kaining Ying 외 arxiv

This paper proposes a large-scale multi-modal dataset for referring motion expression video segmentation, focusing on segmenting and tracking target objects in videos based on language description of objects' motions. Ex…

Referring Video Object SegmentationMulti-Object TrackingVideo SegmentationVideo Captioning

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

2024-12-18 · Cong Wei, Yujie Zhong, Haoxian Tan, Yingsen Zeng 외

Boosted by Multi-modal Large Language Models (MLLMs), text-guided universal segmentation models for the image and video domains have made rapid progress recently. However, these methods are often developed separately for…

Reasoning SegmentationSegmentationUniversal SegmentationVideo Segmentation+2

Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks

2022-01-01 · CVPR 2022 1 · Wenwen Pan, Haonan Shi, Zhou Zhao, Jieming Zhu 외

Audio-Guided video semantic segmentation is a challenging problem in visual analysis and editing, which automatically separates foreground objects from background in a video sequence according to the referring audio …

DecoderDenoisingSegmentationSemantic Segmentation+2