paper-with-me

홈 › Papers

3rd Place Solution for MeViS Track in CVPR 2024 PVUW workshop: Motion Expression guided Video Segmentation

2024-06-07 · Feiyu Pan, Hao Fang, Xiankai Lu

Referring video object segmentation (RVOS) relies on natural language expressions to segment target objects in video, emphasizing modeling dense text-video relations. The current RVOS methods typically use independently pre-trained vision and language models as backbones, resulting in a significant domain gap between video and text. In cross-modal feature interaction, text features are only used as query initialization and do not fully utilize important information in the text. In this work, we propose using frozen pre-trained vision-language models (VLM) as backbones, with a specific emphasis on enhancing cross-modal feature interaction. Firstly, we use frozen convolutional CLIP backbone to generate feature-aligned vision and text features, alleviating the issue of domain gap and reducing training costs. Secondly, we add more cross-modal feature fusion in the pipeline to enhance the utilization of multi-modal information. Furthermore, we propose a novel video query initialization method to generate higher quality video queries. Without bells and whistles, our method achieved 51.5 J&F on the MeViS test set and ranked 3rd place for MeViS Track in CVPR 2024 PVUW workshop: Motion Expression guided Video Segmentation.

📄 PDF Abstract BibTeX arXiv:2406.04842

Code (0)

등록된 구현이 없습니다.

Tasks

Referring Video Object SegmentationSemantic SegmentationVideo Object SegmentationVideo SegmentationVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

1st Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation

2024-06-11 · Mingqi Gao, Jingnan Luo, Jinyu Yang, Jungong Han 외

Motion Expression guided Video Segmentation (MeViS), as an emerging task, poses many new challenges to the field of referring video object segmentation (RVOS). In this technical report, we investigated and validated the …

Referring Video Object SegmentationSegmentationSemantic SegmentationVideo Object Segmentation+2

2nd Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation

2024-06-20 · Bin Cao, Yisi Zhang, Xuanxu Lin, Xingjian He 외

Motion Expression guided Video Segmentation is a challenging task that aims at segmenting objects in the video based on natural language expressions with motion descriptions. Unlike the previous referring video object se…

Instance SegmentationReferring Video Object SegmentationSegmentationSemantic Segmentation+4

The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation

2025-04-07 · Hao Fang, Runmin Cong, Xiankai Lu, Zhiyang Chen 외

Motion expression video segmentation is designed to segment objects in accordance with the input motion expressions. In contrast to the conventional Referring Video Object Segmentation (RVOS), it places emphasis on motio…

Inference OptimizationReferring Video Object SegmentationSegmentationSemantic Segmentation+3

ReferDINO-Plus: 2nd Solution for 4th PVUW MeViS Challenge at CVPR 2025

2025-03-30 · Tianming Liang, Haichao Jiang, Wei-Shi Zheng, Jian-Fang Hu

Referring Video Object Segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This task has attracted increasing attention in the field of computer vision due to its promising …

ObjectReferring Video Object SegmentationSemantic SegmentationVideo Editing+2

Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding

2026-04-28 · Chang Liu, Henghui Ding, Nikhila Ravi, Yunchao Wei 외 arxiv

This report summarizes the objectives, datasets, and top-performing methodologies of the 2026 Pixel-level Video Understanding in the Wild (PVUW) Challenge, hosted at CVPR 2026, which evaluates state-of-the-art models und…

Object Segmentation