paper-with-me

홈 › Papers

Video Jigsaw: Unsupervised Learning of Spatiotemporal Context for Video Action Recognition

2018-08-22 · Unaiza Ahsan, Rishi Madhok, Irfan Essa

We propose a self-supervised learning method to jointly reason about spatial and temporal context for video recognition. Recent self-supervised approaches have used spatial context [9, 34] as well as temporal coherency [32] but a combination of the two requires extensive preprocessing such as tracking objects through millions of video frames [59] or computing optical flow to determine frame regions with high motion [30]. We propose to combine spatial and temporal context in one self-supervised framework without any heavy preprocessing. We divide multiple video frames into grids of patches and train a network to solve jigsaw puzzles on these patches from multiple frames. So the network is trained to correctly identify the position of a patch within a video frame as well as the position of a patch over time. We also propose a novel permutation strategy that outperforms random permutations while significantly reducing computational and memory constraints. We use our trained network for transfer learning tasks such as video activity recognition and demonstrate the strength of our approach on two benchmark video action recognition datasets without using a single frame from these datasets for unsupervised pretraining of our proposed video jigsaw network.

📄 PDF Abstract BibTeX arXiv:1808.07507

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionActivity RecognitionOptical Flow EstimationPositionSelf-Supervised LearningTemporal Action LocalizationTransfer LearningVideo Recognition

Methods 이 논문이 사용한 방법론

Jigsaw Jigsaw is a self-supervision approach that relies on jigsaw-like puzzles as the pretext task in order to learn image representations.

Similar Papers 제목 키워드 기반

Self-Supervised Video Representation Learning with Constrained Spatiotemporal Jigsaw

2021-01-01 · Yuqi Huo, Mingyu Ding, Haoyu Lu, Zhiwu Lu 외

This paper proposes a novel pretext task for self-supervised video representation learning by exploiting spatiotemporal continuity in videos. It is motivated by the fact that videos are spatiotemporal by nature and a rep…

Representation Learning

Using 3D Convolutional Neural Networks to Learn Spatiotemporal Features for Automatic Surgical Gesture Recognition in Video

2019-07-26 · Isabel Funke, Sebastian Bodenstedt, Florian Oehme, Felix von Bechtolsheim 외

Automatically recognizing surgical gestures is a crucial step towards a thorough understanding of surgical skill. Possible areas of application include automatic skill assessment, intra-operative monitoring of critical s…

Gesture RecognitionSurgical Gesture Recognition

Unsupervised Learning of Spatiotemporally Coherent Metrics

2014-12-18 · ICCV 2015 12 · Ross Goroshin, Joan Bruna, Jonathan Tompson, David Eigen 외

Current state-of-the-art classification and detection algorithms rely on supervised training. In this work we study unsupervised feature learning in the context of temporally coherent video data. We focus on feature lear…

General ClassificationMetric Learning

Contrast-Phys: Unsupervised Video-based Remote Physiological Measurement via Spatiotemporal Contrast

2022-08-08 · Zhaodong Sun, Xiaobai Li

Video-based remote physiological measurement utilizes face videos to measure the blood volume change signal, which is also called remote photoplethysmography (rPPG). Supervised methods for rPPG measurements achieve state…

Unsupervised Trajectory Segmentation and Promoting of Multi-Modal Surgical Demonstrations

2018-10-01 · Zhenzhou Shao, Hongfa Zhao, Jiexin Xie, Ying Qu 외

To improve the efficiency of surgical trajectory segmentation for robot learning in robot-assisted minimally invasive surgery, this paper presents a fast unsupervised method using video and kinematic data, followed by a …

Segmentation