paper-with-me

홈 › Papers

Referring to Objects in Videos using Spatio-Temporal Identifying Descriptions

2019-04-08 · WS 2019 6 · Peratham Wiriyathammabhum, Abhinav Shrivastava, Vlad I. Morariu, Larry S. Davis

This paper presents a new task, the grounding of spatio-temporal identifying descriptions in videos. Previous work suggests potential bias in existing datasets and emphasizes the need for a new data creation schema to better model linguistic structure. We introduce a new data collection scheme based on grammatical constraints for surface realization to enable us to investigate the problem of grounding spatio-temporal identifying descriptions in videos. We then propose a two-stream modular attention network that learns and grounds spatio-temporal identifying descriptions based on appearance and motion. We show that motion modules help to ground motion-related words and also help to learn in appearance modules because modular neural networks resolve task interference between modules. Finally, we propose a future challenge and a need for a robust system arising from replacing ground truth visual annotations with automatic video object detector and temporal event localization.

📄 PDF Abstract BibTeX arXiv:1904.03885

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Object Referring in Videos with Language and Human Gaze

2018-01-04 · CVPR 2018 6 · Arun Balajee Vasudevan, Dengxin Dai, Luc van Gool

We investigate the problem of object referring (OR) i.e. to localize a target object in a visual scene coming with a language description. Humans perceive the world more as continued video snippets than as static images,…

ObjectReferring Expression

Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation

2025-05-19 · Tianming Liang, Haichao Jiang, Yuting Yang, Chaolei Tan 외

Referring video object segmentation (RVOS) aims to identify, track and segment the objects in a video based on language descriptions, which has received great attention in recent years. However, existing datasets remain …

Referring Video Object SegmentationSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation

VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation

2026-03-28 · Jihwan Hong, Jaeyoung Do arxiv

Referring Video Object Segmentation (RVOS) aims to segment target objects in videos based on natural language descriptions. However, fixed keyframe-based approaches that couple a vision language model with a separate pro…

Referring Video Object Segmentation

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension

2026-07-02 · Shunya Kato, Taiki Miyanishi, Shuhei Kurita, Mahiro Ukai 외 arxiv

Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this context, Video Referring Expression Comprehen…

Referring Expression

SVAC: Scaling Is All You Need For Referring Video Object Segmentation

2025-09-28 · Li Zhang, Haoxiang Gao, Zhihao Zhang, Luoxiao Huang 외 arxiv

Referring Video Object Segmentation (RVOS) aims to segment target objects in video sequences based on natural language descriptions. While recent advances in Multi-modal Large Language Models (MLLMs) have improved RVOS p…

Referring Video Object Segmentation