paper-with-me

Papers

Unsupervised Video Object Segmentation with Joint Hotspot Tracking

2020-08-01 · ECCV 2020 8 · Lu Zhang, Jianming Zhang, Zhe Lin, Radomír Měch, Huchuan Lu, You He

Object tracking is a well-studied problem in computer vision while identifying salient spots of objects in a video is a less explored direction in the literature. Video eye gaze estimation methods aim to tackle a related task but salient spots in those methods are not bounded by objects and tend to produce very scattered, unstable predictions due to the noisy ground truth data. We reformulate the problem of detecting and tracking of salient object spots as a new task called object hotspot tracking. In this paper, we propose to tackle this task jointly with unsupervised video object segmentation, in real-time, with a unified framework to exploit the synergy between the two. Specifically, we propose a Weighted Correlation Siamese Network (WCS-Net) which employs a Weighted Correlation Block (WCB) for encoding the pixel-wise correspondence between a template frame and the search frame. In addition, WCB takes the initial mask / hotspot as guidance to enhance the influence of salient regions for robust tracking. Our system can operate online during inference and jointly produce the object mask and hotspot track-lets at 33 FPS. Experimental results validate the effectiveness of our network design, and show the benefits of jointly solving the hotspot tracking and object segmentation problems. In particular, our method performs favorably against state-of-the-art video eye gaze models in object hotspot tracking, and outperforms existing methods on three benchmark datasets for unsupervised video object segmentation.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Gaze EstimationObjectObject TrackingSegmentationSemantic SegmentationUnsupervised Video Object SegmentationVideo Object SegmentationVideo Semantic Segmentation

Similar Papers 제목 키워드 기반

Grounded Human-Object Interaction Hotspots from Video (Extended Abstract)

2019-06-03 · Tushar Nagarajan, Christoph Feichtenhofer, Kristen Grauman

Learning how to interact with objects is an important step towards embodied visual intelligence, but existing techniques suffer from heavy supervision or sensing requirements. We propose an approach to learn human-object…

Human-Object Interaction DetectionObjectSemantic Segmentation

Grounded Human-Object Interaction Hotspots from Video

2018-12-11 · ICCV 2019 10 · Tushar Nagarajan, Christoph Feichtenhofer, Kristen Grauman

Learning how to interact with objects is an important step towards embodied visual intelligence, but existing techniques suffer from heavy supervision or sensing requirements. We propose an approach to learn human-object…

Human-Object Interaction DetectionObjectObject RecognitionSemantic Segmentation+1

Joint Hand Motion and Interaction Hotspots Prediction from Egocentric Videos

2022-04-04 · CVPR 2022 1 · Shaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong Wang

We propose to forecast future hand-object interactions given an egocentric video. Instead of predicting action labels or pixels, we directly predict the hand motion trajectory and the future contact points on the next ac…

Object

Scene-Centric Unsupervised Video Panoptic Segmentation

2026-06-03 · Christoph Reich, Oliver Hahn, Nikita Araslanov, Laura Leal-Taixé 외 arxiv

Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervised VPS, omitting any …

Video Panoptic SegmentationScene UnderstandingVideo SegmentationImage Segmentation

Forecasting Human-Object Interaction: Joint Prediction of Motor Attention and Actions in First Person Video

2019-11-25 · ECCV 2020 8 · Miao Liu, Siyu Tang, Yin Li, James Rehg

We address the challenging task of anticipating human-object interaction in first person videos. Most existing methods ignore how the camera wearer interacts with the objects, or simply consider body motion as a separate…

Action AnticipationHuman-Object Interaction Detection