paper-with-me

홈 › Papers

YouMVOS: An Actor-Centric Multi-Shot Video Object Segmentation Dataset

2022-01-01 · CVPR 2022 1 · Donglai Wei, Siddhant Kharbanda, Sarthak Arora, Roshan Roy, Nishant Jain, Akash Palrecha, Tanav Shah, Shray Mathur, Ritik Mathur, Abhijay Kemkar, Anirudh Chakravarthy, Zudi Lin, Won-Dong Jang, Yansong Tang, Song Bai, James Tompkin, Philip H.S. Torr, Hanspeter Pfister

Many video understanding tasks require analyzing multi-shot videos, but existing datasets for video object segmentation (VOS) only consider single-shot videos. To address this challenge, we collected a new dataset---YouMVOS---of 200 popular YouTube videos spanning ten genres, where each video is on average five minutes long and with 75 shots. We selected recurring actors and annotated 431K segmentation masks at a frame rate of six, exceeding previous datasets in average video duration, object variation, and narrative structure complexity. We incorporated good practices of model architecture design, memory management, and multi-shot tracking into an existing video segmentation method to build competitive baseline methods. Through error analysis, we found that these baselines still fail to cope with cross-shot appearance variation on our YouMVOS dataset. Thus, our dataset poses new challenges in multi-shot segmentation towards better video analysis. Data, code, and pre-trained models are available at https://donglaiw.github.io/proj/youMVOS

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ManagementSegmentationSemantic SegmentationVideo Object SegmentationVideo SegmentationVideo Semantic SegmentationVideo Understanding

Similar Papers 제목 키워드 기반

Segment Anything Across Shots: A Method and Benchmark

2025-11-17 · Hengrui Hu, Kaining Ying, Henghui Ding arxiv

This work focuses on multi-shot semi-supervised video object segmentation (MVOS), which aims at segmenting the target object indicated by an initial mask throughout a video with multiple shots. The existing VOS methods m…

Semi-Supervised Video Object SegmentationData Augmentation

Object-Shot Enhanced Grounding Network for Egocentric Video

2025-05-07 · CVPR 2025 1 · Yisen Feng, Haoyu Zhang, Meng Liu, Weili Guan 외

Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the distributional differences between egocentr…

Video Grounding

Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion

2026-05-25 · Dmytro Klepachevskyi, Alexander Wong, Sirisha Rambhatla, Yuhao Chen arxiv

Object re-identification (ReID) in egocentric kitchen videos is challenging due to rapid viewpoint changes, frequent occlusions, cluttered scenes, and large intra-class appearance variations. Objects may leave and re-ent…

WildActor: Unconstrained Identity-Preserving Video Generation

2026-02-28 · Qin Guo, Tianyu Yang, Xuanhua He, Fei Shen 외 arxiv

Production-ready human video generation requires digital actors to maintain strictly consistent full-body identities across dynamic shots, viewpoints and motions, a setting that remains challenging for existing methods. …

Video Generation

Storyline Representation of Egocentric Videos With an Applications to Story-Based Search

2015-12-01 · ICCV 2015 12 · Bo Xiong, Gunhee Kim, Leonid Sigal

Egocentric videos are a valuable source of information as a daily log of our lives. However, large fraction of egocentric video content is typically irrelevant and boring to re-watch. It is an agonizing task, for example…