paper-with-me

홈 › Papers

UNINEXT-Cutie: The 1st Solution for LSVOS Challenge RVOS Track

2024-08-19 · Hao Fang, Feiyu Pan, Xiankai Lu, Wei zhang, Runmin Cong

Referring video object segmentation (RVOS) relies on natural language expressions to segment target objects in video. In this year, LSVOS Challenge RVOS Track replaced the origin YouTube-RVOS benchmark with MeViS. MeViS focuses on referring the target object in a video through its motion descriptions instead of static attributes, posing a greater challenge to RVOS task. In this work, we integrate strengths of that leading RVOS and VOS models to build up a simple and effective pipeline for RVOS. Firstly, We finetune the state-of-the-art RVOS model to obtain mask sequences that are correlated with language descriptions. Secondly, based on a reliable and high-quality key frames, we leverage VOS model to enhance the quality and temporal consistency of the mask results. Finally, we further improve the performance of the RVOS model using semi-supervised learning. Our solution achieved 62.57 J&F on the MeViS test set and ranked 1st place for 6th LSVOS Challenge RVOS Track.

📄 PDF Abstract BibTeX arXiv:2408.10129

Code (0)

등록된 구현이 없습니다.

Tasks

Referring Video Object SegmentationSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
VOS VOS is a type of video object segmentation model consisting of two network components. The target appearance model consists of a light-weight module, which is learned during…

Similar Papers 제목 키워드 기반

Enriched Feature Representation and Motion Prediction Module for MOSEv2 Track of 7th LSVOS Challenge: 3rd Place Solution

2025-09-19 · Chang Soo Lim, Joonyoung Moon, Donghyeon Cho arxiv

Video object segmentation (VOS) is a challenging task with wide applications such as video editing and autonomous driving. While Cutie provides strong query-based segmentation and SAM2 offers enriched representations via…

Video Object SegmentationAutonomous Driving

LSVOS Challenge 3rd Place Report: SAM2 and Cutie based VOS

2024-08-20 · Xinyu Liu, Jing Zhang, Kexin Zhang, Xu Liu 외

Video Object Segmentation (VOS) presents several challenges, including object occlusion and fragmentation, the dis-appearance and re-appearance of objects, and tracking specific objects within crowded scenes. In this wor…

Instance SegmentationObjectSegmentationSemantic Segmentation+3

The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA

2025-09-21 · Quanzhu Niu, Dengxian Gong, Shihao Chen, Tao Zhang 외 arxiv

Referring video object segmentation (RVOS) requires segmenting and tracking objects in videos conditioned on natural-language expressions, demanding fine-grained understanding of both appearance and motion. Building on S…

Referring Video Object SegmentationVideo Segmentation

The 2nd Solution for LSVOS Challenge RVOS Track: Spatial-temporal Refinement for Consistent Semantic Segmentation

2024-08-22 · Tuyen Tran

Referring Video Object Segmentation (RVOS) is a challenging task due to its requirement for temporal understanding. Due to the obstacle of computational complexity, many state-of-the-art models are trained on short time …

Referring Video Object SegmentationSegmentationSemantic SegmentationVideo Object Segmentation+1

Enhancing Sa2VA for Referent Video Object Segmentation: 2nd Solution for 7th LSVOS RVOS Track

2025-09-19 · Ran Hong, Feng Lu, Leilei Cao, An Yan 외 arxiv

Referential Video Object Segmentation (RVOS) aims to segment all objects in a video that match a given natural language description, bridging the gap between vision and language understanding. Recent work, such as Sa2VA,…

Video Object SegmentationVideo Segmentation