paper-with-me

Papers

MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions

2023-08-16 · ICCV 2023 1 · Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, Chen Change Loy

This paper strives for motion expressions guided video segmentation, which focuses on segmenting objects in video content based on a sentence describing the motion of the objects. Existing referring video object datasets typically focus on salient objects and use language expressions that contain excessive static attributes that could potentially enable the target object to be identified in a single frame. These datasets downplay the importance of motion in video content for language-guided video object segmentation. To investigate the feasibility of using motion expressions to ground and segment objects in videos, we propose a large-scale dataset called MeViS, which contains numerous motion expressions to indicate target objects in complex environments. We benchmarked 5 existing referring video object segmentation (RVOS) methods and conducted a comprehensive comparison on the MeViS dataset. The results show that current RVOS methods cannot effectively address motion expression-guided video segmentation. We further analyze the challenges and propose a baseline approach for the proposed MeViS dataset. The goal of our benchmark is to provide a platform that enables the development of effective language-guided video segmentation algorithms that leverage motion expressions as a primary cue for object segmentation in complex video scenes. The proposed MeViS dataset has been released at https://henghuiding.github.io/MeViS.

📄 PDF Abstract BibTeX arXiv:2308.08544

Code (1)

henghuiding/MeViS 공식 구현 pytorch

Tasks

Motion Expressions Guided Video SegmentationObjectReferring Video Object SegmentationSegmentationSemantic SegmentationSentenceVideo Object SegmentationVideo SegmentationVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation

2025-12-11 · Henghui Ding, Chang Liu, Shuting He, Kaining Ying 외 arxiv

This paper proposes a large-scale multi-modal dataset for referring motion expression video segmentation, focusing on segmenting and tracking target objects in videos based on language description of objects' motions. Ex…

Referring Video Object SegmentationMulti-Object TrackingVideo SegmentationVideo Captioning

1st Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation

2024-06-11 · Mingqi Gao, Jingnan Luo, Jinyu Yang, Jungong Han 외

Motion Expression guided Video Segmentation (MeViS), as an emerging task, poses many new challenges to the field of referring video object segmentation (RVOS). In this technical report, we investigated and validated the …

Referring Video Object SegmentationSegmentationSemantic SegmentationVideo Object Segmentation+2

LSVOS Challenge Report: Large-scale Complex and Long Video Object Segmentation

2024-09-09 · Henghui Ding, Lingyi Hong, Chang Liu, Ning Xu 외

Despite the promising performance of current video segmentation models on existing benchmarks, these models still struggle with complex scenes. In this paper, we introduce the 6th Large-scale Video Object Segmentation (L…

ObjectReferring Video Object SegmentationSegmentationSemantic Segmentation+3

4th PVUW MeViS 3rd Place Report: Sa2VA

2025-04-01 · Haobo Yuan, Tao Zhang, Xiangtai Li, Lu Qi 외

Referring video object segmentation (RVOS) is a challenging task that requires the model to segment the object in a video given the language description. MeViS is a recently proposed dataset that contains motion expressi…

Language ModelingLanguage ModellingLarge Language ModelReferring Expression+4

The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation

2025-04-07 · Hao Fang, Runmin Cong, Xiankai Lu, Zhiyang Chen 외

Motion expression video segmentation is designed to segment objects in accordance with the input motion expressions. In contrast to the conventional Referring Video Object Segmentation (RVOS), it places emphasis on motio…

Inference OptimizationReferring Video Object SegmentationSegmentationSemantic Segmentation+3