paper-with-me

홈 › Papers

VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning

2025-11-20 · Zishan Xu, Yifu Guo, Yuquan Lu, Fengyu Yang, Junxin Li arxiv

Traditional video reasoning segmentation methods rely on supervised fine-tuning, which limits generalization to out-of-distribution scenarios and lacks explicit reasoning. To address this, we propose \textbf{VideoSeg-R1}, the first framework to introduce reinforcement learning into video reasoning segmentation. It adopts a decoupled architecture that formulates the task as joint referring image segmentation and video mask propagation. It comprises three stages: (1) A hierarchical text-guided frame sampler to emulate human attention; (2) A reasoning model that produces spatial cues along with explicit reasoning chains; and (3) A segmentation-propagation stage using SAM2 and XMem. A task difficulty-aware mechanism adaptively controls reasoning length for better efficiency and accuracy. Extensive evaluations on multiple benchmarks demonstrate that VideoSeg-R1 achieves state-of-the-art performance in complex video reasoning and segmentation tasks. The code will be publicly available at https://github.com/euyis1019/VideoSeg-R1.

📄 PDF Abstract BibTeX arXiv:2511.16077

Code (0)

등록된 구현이 없습니다.

Tasks

Video Object SegmentationReinforcement LearningImage Segmentation

Similar Papers 제목 키워드 기반

VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation

2026-06-05 · Ming Dai, Sen Yang, Boqiang Duan, Boyuan Tong 외 arxiv

Reasoning Video Object Segmentation (RVOS) demands a sophisticated integration of temporal dynamics, spatial details, and linguistic reasoning to achieve precise pixel-level localization. Existing methods are limited to …

Video Object SegmentationReinforcement Learning

Flow-based Video Segmentation for Human Head and Shoulders

2021-04-20 · Zijian Kuang, Xinran Tie

Video segmentation for the human head and shoulders is essential in creating elegant media for videoconferencing and virtual reality applications. The main challenge is to process high-quality background subtraction in a…

DecoderImage MattingImage SegmentationOptical Flow Estimation+5

ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning

2025-12-02 · Yifan Li, Yingda Yin, Lingting Zhu, Weikai Chen 외 arxiv

Reasoning-centric video object segmentation is an inherently complex task: the query often refers to dynamics, causality, and temporal interactions, rather than static appearances. Yet existing solutions generally collap…

Video Object SegmentationReinforcement LearningVideo Segmentation

FeVOS: Foresight Expression Video Object Segmentation

2026-06-24 · Kehan Lan, Kaining Ying, Henghui Ding arxiv

Existing Referring Video Object Segmentation tasks focus on referring expressions describing events, actions or appearances of relevant objects within the observed frames, lacking evaluation in scenarios that require pre…

Referring Video Object SegmentationReinforcement Learning

ROLL: Visual Self-Supervised Reinforcement Learning with Object Reasoning

2020-11-13 · YuFei Wang, Gautham Narayan Narasimhan, Xingyu Lin, Brian Okorn 외

Current image-based reinforcement learning (RL) algorithms typically operate on the whole image without performing object-level reasoning. This leads to inefficient goal sampling and ineffective reward functions. In this…

Multi-Goal Reinforcement LearningObjectreinforcement-learningReinforcement Learning+1