paper-with-me

홈 › Papers

FeVOS: Foresight Expression Video Object Segmentation

2026-06-24 · Kehan Lan, Kaining Ying, Henghui Ding arxiv

Existing Referring Video Object Segmentation tasks focus on referring expressions describing events, actions or appearances of relevant objects within the observed frames, lacking evaluation in scenarios that require pre-decisive spatio-temporal reasoning, thereby limiting their applicability. To address this, we propose Foresight Expression Video Object Segmentation, a task that queries future events in upcoming video segments and requires masks of the objects in the observed frames as visual answers. For example, in ego-centric scenes, the question "What tool will be used?" demands reasoning over spatio-temporal cues to predict the masks of the next tool to be used, which helps with the understanding of future actions and decisions. To support this task, we introduce FeVOS, a dataset with 968 video clips, 14,525 foresight expressions, and 2,904 chain-of-thought annotations to provide explicit and interpretable reasoning steps. We further develop FeVOS-R1, an MLLM-based model trained on our dataset via a two-stage pipeline of supervised fine-tuning and reinforcement learning. FeVOS-R1 not only achieves state-of-the-art performance on FeVOS, but also demonstrates strong generalization to existing RVOS benchmarks. We hope this work can inspire more research on predictive reasoning in video perception.

📄 PDF Abstract BibTeX arXiv:2606.25585

Code (0)

등록된 구현이 없습니다.

Tasks

Referring Video Object SegmentationReinforcement Learning

Similar Papers 제목 키워드 기반

ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos

2026-01-08 · Rustin Soraki, Homanga Bharadhwaj, Ali Farhadi, Roozbeh Mottaghi arxiv

Humans can effortlessly anticipate how objects might move or change through interaction--imagining a cup being lifted, a knife slicing, or a lid being closed. We aim to endow computational systems with a similar ability …

3D Pose Estimation

MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions

2023-08-16 · ICCV 2023 1 · Henghui Ding, Chang Liu, Shuting He, Xudong Jiang 외

This paper strives for motion expressions guided video segmentation, which focuses on segmenting objects in video content based on a sentence describing the motion of the objects. Existing referring video object datasets…

Motion Expressions Guided Video SegmentationObjectReferring Video Object SegmentationSegmentation+5

MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation

2025-12-11 · Henghui Ding, Chang Liu, Shuting He, Kaining Ying 외 arxiv

This paper proposes a large-scale multi-modal dataset for referring motion expression video segmentation, focusing on segmenting and tracking target objects in videos based on language description of objects' motions. Ex…

Referring Video Object SegmentationMulti-Object TrackingVideo SegmentationVideo Captioning

MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction

2026-07-17 · Homanga Bharadhwaj, Yash Jangir arxiv

Humans can infer how objects are likely to move from passive observation: a cup may be lifted, a drawer may slide, and a lid may rotate shut. Such predictions expose the physical consequences of interaction needed to act…

Motion ForecastingVideo Prediction

Decoupled Motion Expression Video Segmentation

2025-01-01 · CVPR 2025 1 · Hao Fang, Runmin Cong, Xiankai Lu, Xiaofei Zhou 외

Motion expression video segmentation aims to segment objects based on input motion descriptions. Compared with traditional referring video object segmentation, it focuses on motion and multi-object expressions and is…

Instance SegmentationReferring Video Object SegmentationSegmentationSemantic Segmentation+4