Exploiting Motion Information from Unlabeled Videos for Static Image Action Recognition
Static image action recognition, which aims to recognize action based on a single image, usually relies on expensive human labeling effort such as adequate labeled action images and large-scale labeled image dataset. In contrast, abundant unlabeled videos can be economically obtained. Therefore, several works have explored using unlabeled videos to facilitate image action recognition, which can be categorized into the following two groups: (a) enhance visual representations of action images with a designed proxy task on unlabeled videos, which falls into the scope of self-supervised learning; (b) generate auxiliary representations for action images with the generator learned from unlabeled videos. In this paper, we integrate the above two strategies in a unified framework, which consists of Visual Representation Enhancement (VRE) module and Motion Representation Augmentation (MRA) module. Specifically, the VRE module includes a proxy task which imposes pseudo motion label constraint and temporal coherence constraint on unlabeled videos, while the MRA module could predict the motion information of a static action image by exploiting unlabeled videos. We demonstrate the superiority of our framework based on four benchmark human action datasets with limited labeled data.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionSelf-Supervised LearningSimilar Papers 제목 키워드 기반
MOD-UV: Learning Mobile Object Detectors from Unlabeled Videos
Embodied agents must detect and localize objects of interest, e.g. traffic participants for self-driving cars. Supervision in the form of bounding boxes for this task is extremely expensive. As such, prior work has looke…
Motion SegmentationObjectobject-detectionObject Detection+5MotionZero:Exploiting Motion Priors for Zero-shot Text-to-Video Generation
Zero-shot Text-to-Video synthesis generates videos based on prompts without any videos. Without motion information from videos, motion priors implied in prompts are vital guidance. For example, the prompt "airplane landi…
DisentanglementText-to-Video GenerationVideo GenerationZero-shot Text-to-Video GenerationIm2Flow: Motion Hallucination from Static Images for Action Recognition
Existing methods to recognize actions in static images take the images at their face value, learning the appearances---objects, scenes, and body poses---that distinguish each action class. However, such models are depriv…
Action RecognitionActivity RecognitionDecoderHallucination+2Joint Unsupervised Learning of Optical Flow and Depth by Watching Stereo Videos
Learning depth and optical flow via deep neural networks by watching videos has made significant progress recently. In this paper, we jointly solve the two tasks by exploiting the underlying geometric rules within stereo…
Motion EstimationOptical Flow EstimationPonymation: Learning 3D Animal Motions from Unlabeled Online Videos
We introduce Ponymation, a new method for learning a generative model of articulated 3D animal motions from raw, unlabeled online videos. Unlike existing approaches for motion synthesis, our model does not require any po…
Motion Synthesis