paper-with-me

Papers

EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations

2022-09-26 · Ahmad Darkhalil, Dandan Shan, Bin Zhu, Jian Ma, Amlan Kar, Richard Higgins, Sanja Fidler, David Fouhey, Dima Damen

We introduce VISOR, a new dataset of pixel annotations and a benchmark suite for segmenting hands and active objects in egocentric video. VISOR annotates videos from EPIC-KITCHENS, which comes with a new set of challenges not encountered in current video segmentation datasets. Specifically, we need to ensure both short- and long-term consistency of pixel-level annotations as objects undergo transformative interactions, e.g. an onion is peeled, diced and cooked - where we aim to obtain accurate pixel-level annotations of the peel, onion pieces, chopping board, knife, pan, as well as the acting hands. VISOR introduces an annotation pipeline, AI-powered in parts, for scalability and quality. In total, we publicly release 272K manual semantic masks of 257 object classes, 9.9M interpolated dense masks, 67K hand-object relations, covering 36 hours of 179 untrimmed videos. Along with the annotations, we introduce three challenges in video object segmentation, interaction understanding and long-term reasoning. For data, code and leaderboards: http://epic-kitchens.github.io/VISOR

📄 PDF Abstract BibTeX arXiv:2209.13064

Code (3)

epic-kitchens/visor-hos 공식 구현 pytorch
epic-kitchens/visor-vos 공식 구현 pytorch
epic-kitchens/visor-wdtcf 공식 구현

Tasks

ObjectSegmentationSemantic SegmentationVideo Object SegmentationVideo SegmentationVideo Semantic Segmentation

Similar Papers 제목 키워드 기반

EPIC Fields: Marrying 3D Geometry and Video Understanding

2023-06-14 · NeurIPS 2023 11 · Vadim Tschernezki, Ahmad Darkhalil, Zhifan Zhu, David Fouhey 외

Neural rendering is fuelling a unification of learning, 3D geometry and video understanding that has been waiting for more than two decades. Progress, however, is still hampered by a lack of suitable datasets and benchma…

3D geometryNeural RenderingVideo Understanding

Team I2R-VI-FF Technical Report on EPIC-KITCHENS VISOR Hand Object Segmentation Challenge 2023

2023-10-31 · Fen Fang, Yi Cheng, Ying Sun, Qianli Xu

In this report, we present our approach to the EPIC-KITCHENS VISOR Hand Object Segmentation Challenge, which focuses on the estimation of the relation between the hands and the objects given a single frame as input. The …

Hand SegmentationObjectSegmentationSemantic Segmentation

Forecasting Action through Contact Representations from First Person Video

2021-02-01 · Eadom Dessalene, Chinmaya Devaraj, Michael Maynord, Cornelia Fermuller 외

Human actions involving hand manipulations are structured according to the making and breaking of hand-object contact, and human visual understanding of action is reliant on anticipation of contact as is demonstrated by …

Action AnticipationObject

Anticipative Video Transformer

2021-06-03 · ICCV 2021 10 · Rohit Girdhar, Kristen Grauman

We propose Anticipative Video Transformer (AVT), an end-to-end attention-based video modeling architecture that attends to the previously observed video in order to anticipate future actions. We train the model jointly t…

Action Anticipation

Scaling Egocentric Vision: The EPIC-KITCHENS Dataset

2018-04-08 · ECCV 2018 9 · Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler 외

First-person vision is gaining interest as it offers a unique viewpoint on people's interaction with objects, their attention, and even intention. However, progress in this challenging domain has been relatively slow due…

Action Anticipation