Geometry Guided Convolutional Neural Networks for Self-Supervised Video Representation Learning
It is often laborious and costly to manually annotate videos for training high-quality video recognition models, so there has been some work and interest in exploring alternative, cheap, and yet often noisy and indirect, training signals for learning the video representations. However, these signals are still coarse, supplying supervision at the whole video frame level, and subtle, sometimes enforcing the learning agent to solve problems that are even hard for humans. In this paper, we instead explore geometry, a grand new type of auxiliary supervision for the self-supervised learning of video representations. In particular, we extract pixel-wise geometry information as flow fields and disparity maps from synthetic imagery and real 3D movies. Although the geometry and high-level semantics are seemingly distant topics, surprisingly, we find that the convolutional neural networks pre-trained by the geometry cues can be effectively adapted to semantic video understanding tasks. In addition, we also find that a progressive training strategy can foster a better neural network for the video recognition task than blindly pooling the distinct sources of geometry cues together. Extensive results on video dynamic scene recognition and action recognition tasks show that our geometry guided networks significantly outperform the competing methods that are trained with other types of labeling-free supervision signals.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionRepresentation LearningScene RecognitionSelf-Supervised LearningTemporal Action LocalizationVideo RecognitionVideo UnderstandingSimilar Papers 제목 키워드 기반
Learning a Spatio-Temporal Embedding for Video Instance Segmentation
We present a novel embedding approach for video instance segmentation. Our method learns a spatio-temporal embedding integrating cues from appearance, motion, and geometry; a 3D causal convolutional network models motion…
Instance SegmentationSemantic SegmentationVideo Instance SegmentationSelf-supervised Learning of Occlusion Aware Flow Guided 3D Geometry Perception with Adaptive Cross Weighted Loss from Monocular Videos
Self-supervised deep learning-based 3D scene understanding methods can overcome the difficulty of acquiring the densely labeled ground-truth and have made a lot of advances. However, occlusions and moving objects are sti…
3D geometry3D Geometry PerceptionCamera Pose EstimationOptical Flow Estimation+3Normal-guided Garment UV Prediction for Human Re-texturing
Clothes undergo complex geometric deformations, which lead to appearance changes. To edit human videos in a physically plausible way, a texture map must take into account not only the garment transformation induced by th…
3D ReconstructionPredictionNeural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion
Self-supervised learning has emerged as a powerful tool for depth and ego-motion estimation, leading to state-of-the-art results on benchmark datasets. However, one significant limitation shared by current methods is the…
Depth EstimationMotion EstimationSelf-Supervised LearningVisual OdometrySelf-supervised 3D Representation Learning of Dressed Humans from Social Media Videos
A key challenge of learning a visual representation for the 3D high fidelity geometry of dressed humans lies in the limited availability of the ground truth data (e.g., 3D scanned models), which results in the performanc…
3D Human ReconstructionDepth EstimationRepresentation LearningSelf-Supervised Learning