paper-with-me

Papers

Anticipating Visual Representations from Unlabeled Video

2015-04-29 · CVPR 2016 6 · Carl Vondrick, Hamed Pirsiavash, Antonio Torralba

Anticipating actions and objects before they start or appear is a difficult problem in computer vision with several real-world applications. This task is challenging partly because it requires leveraging extensive knowledge of the world that is difficult to write down. We believe that a promising resource for efficiently learning this knowledge is through readily available unlabeled video. We present a framework that capitalizes on temporal structure in unlabeled video to learn to anticipate human actions and objects. The key idea behind our approach is that we can train deep networks to predict the visual representation of images in the future. Visual representations are a promising prediction target because they encode images at a higher semantic level than pixels yet are automatic to compute. We then apply recognition algorithms on our predicted representation to anticipate objects and actions. We experimentally validate this idea on two datasets, anticipating actions one second in the future and objects five seconds in the future.

📄 PDF Abstract BibTeX arXiv:1504.08023

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SoundNet: Learning Sound Representations from Unlabeled Video

2016-10-27 · NeurIPS 2016 12 · Yusuf Aytar, Carl Vondrick, Antonio Torralba

We learn rich natural sound representations by capitalizing on large amounts of unlabeled sound data collected in the wild. We leverage the natural synchronization between vision and sound to learn an acoustic representa…

General Classification

Local Frequency Domain Transformer Networks for Video Prediction

2021-05-10 · Hafez Farazi, Jan Nogga, Sven Behnke

Video prediction is commonly referred to as forecasting future frames of a video sequence provided several past frames thereof. It remains a challenging domain as visual scenes evolve according to complex underlying dyna…

Motion SegmentationPredictionVideo Prediction

Exploiting Motion Information from Unlabeled Videos for Static Image Action Recognition

2019-12-01 · Yiyi Zhang, Li Niu, Ziqi Pan, Meichao Luo 외

Static image action recognition, which aims to recognize action based on a single image, usually relies on expensive human labeling effort such as adequate labeled action images and large-scale labeled image dataset. In …

Action RecognitionSelf-Supervised Learning

Anticipating Daily Intention using On-Wrist Motion Triggered Sensing

2017-10-20 · ICCV 2017 10 · Tz-Ying Wu, Ting-An Chien, Cheng-Sheng Chan, Chan-Wei Hu 외

Anticipating human intention by observing one's actions has many applications. For instance, picking up a cellphone, then a charger (actions) implies that one wants to charge the cellphone (intention). By anticipating th…

Slow and steady feature analysis: higher order temporal coherence in video

2015-06-15 · CVPR 2016 6 · Dinesh Jayaraman, Kristen Grauman

How can unlabeled video augment visual learning? Existing methods perform "slow" feature analysis, encouraging the representations of temporally close frames to exhibit only small differences. While this standard approac…

Action RecognitionTemporal Action Localization