paper-with-me

홈 › Papers

Temporal Grounding as a Learning Signal for Referring Video Object Segmentation

2025-08-16 · Seunghun Lee, Jiwan Seo, Jeonghoon Kim, Sungho Moon, Siwon Kim, Haeun Yun, Hyogyeong Jeon, Wonhyeok Choi, Jaehoon Jeong, Zane Durante, Sang Hyun Park, Sunghoon Im arxiv

Referring Video Object Segmentation (RVOS) aims to segment and track objects in videos based on natural language expressions, requiring precise alignment between visual content and textual queries. However, existing methods often suffer from semantic misalignment, largely due to indiscriminate frame sampling and supervision of all visible objects during training -- regardless of their actual relevance to the expression. We identify the core problem as the absence of an explicit temporal learning signal in conventional training paradigms. To address this, we introduce MeViS-M, a dataset built upon the challenging MeViS benchmark, where we manually annotate temporal spans when each object is referred to by the expression. These annotations provide a direct, semantically grounded supervision signal that was previously missing. To leverage this signal, we propose Temporally Grounded Learning (TGL), a novel learning framework that directly incorporates temporal grounding into the training process. Within this frame- work, we introduce two key strategies. First, Moment-guided Dual-path Propagation (MDP) improves both grounding and tracking by decoupling language-guided segmentation for relevant moments from language-agnostic propagation for others. Second, Object-level Selective Supervision (OSS) supervises only the objects temporally aligned with the expression in each training clip, thereby reducing semantic noise and reinforcing language-conditioned learning. Extensive experiments demonstrate that our TGL framework effectively leverages temporal signal to establish a new state-of-the-art on the challenging MeViS benchmark. We will make our code and the MeViS-M dataset publicly available.

📄 PDF Abstract BibTeX arXiv:2508.11955

Code (0)

등록된 구현이 없습니다.

Tasks

Referring Video Object Segmentation

Similar Papers 제목 키워드 기반

STORM: End-to-End Referring Multi-Object Tracking in Videos

2026-04-12 · Zijia Lu, Jingru Yi, Jue Wang, Yuxiao Chen 외 arxiv

Referring multi-object tracking (RMOT) is a task of associating all the objects in a video that semantically match with given textual queries or referring expressions. Existing RMOT approaches decompose object grounding …

Multi-Object Tracking

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

2025-01-24 · Tianming Liang, Kun-Yu Lin, Chaolei Tan, JianGuo Zhang 외

Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. Despite notable progress in recent years, current RVOS models remain struggle to handle complicate…

DecoderObjectReferring Expression SegmentationReferring Video Object Segmentation+4

SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models

2025-05-24 · Ye Sun, Hao Zhang, Henghui Ding, Tiehua Zhang 외

Achieving fine-grained spatio-temporal understanding in videos remains a major challenge for current Video Large Multimodal Models (Video LMMs). Addressing this challenge requires mastering two core capabilities: video r…

BenchmarkingVideo Grounding

Object-centric Video Question Answering with Visual Grounding and Referring

2025-07-25 · Haochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai 외 arxiv

Video Large Language Models (VideoLLMs) have recently demonstrated remarkable progress in general video understanding. However, existing models primarily focus on high-level comprehension and are limited to text-only res…

Video Question AnsweringObject SegmentationVisual Grounding

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension

2026-07-02 · Shunya Kato, Taiki Miyanishi, Shuhei Kurita, Mahiro Ukai 외 arxiv

Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this context, Video Referring Expression Comprehen…

Referring Expression