paper-with-me

Papers

Predictive Coding Networks Meet Action Recognition

2019-10-22 · Xia Huang, Hossein Mousavi, Gemma Roig

Action recognition is a key problem in computer vision that labels videos with a set of predefined actions. Capturing both, semantic content and motion, along the video frames is key to achieve high accuracy performance on this task. Most of the state-of-the-art methods rely on RGB frames for extracting the semantics and pre-computed optical flow fields as a motion cue. Then, both are combined using deep neural networks. Yet, it has been argued that such models are not able to leverage the motion information extracted from the optical flow, but instead the optical flow allows for better recognition of people and objects in the video. This urges the need to explore different cues or models that can extract motion in a more informative fashion. To tackle this issue, we propose to explore the predictive coding network, so called PredNet, a recurrent neural network that propagates predictive coding errors across layers and time steps. We analyze whether PredNet can better capture motions in videos by estimating over time the representations extracted from pre-trained networks for action recognition. In this way, the model only relies on the video frames, and does not need pre-processed optical flows as input. We report the effectiveness of our proposed model on UCF101 and HMDB51 datasets.

📄 PDF Abstract BibTeX arXiv:1910.10056

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionOptical Flow Estimation

Similar Papers 제목 키워드 기반

Memory-augmented Dense Predictive Coding for Video Representation Learning

2020-08-03 · ECCV 2020 8 · Tengda Han, Weidi Xie, Andrew Zisserman

The objective of this paper is self-supervised learning from video, in particular for representations for action recognition. We make the following contributions: (i) We propose a new architecture and learning framework …

Action ClassificationAction RecognitionOptical Flow EstimationRepresentation Learning+3

Minimal Feature Analysis for Isolated Digit Recognition for varying encoding rates in noisy environments

2022-08-27 · Muskan Garg, Naveen Aggarwal

This research work is about recent development made in speech recognition. In this research work, analysis of isolated digit recognition in the presence of different bit rates and at different noise levels has been perfo…

speech-recognitionSpeech Recognition

MuPNet: Multi-modal Predictive Coding Network for Place Recognition by Unsupervised Learning of Joint Visuo-Tactile Latent Representations

2019-09-16 · Oliver Struckmeier, Kshitij Tiwari, Shirin Dora, Martin J. Pearson 외

Extracting and binding salient information from different sensory modalities to determine common features in the environment is a significant challenge in robotics. Here we present MuPNet (Multi-modal Predictive Coding N…

Video Representation Learning by Dense Predictive Coding

2019-09-10 · Tengda Han, Weidi Xie, Andrew Zisserman

The objective of this paper is self-supervised learning of spatio-temporal embeddings from video, suitable for human action recognition. We make three contributions: First, we introduce the Dense Predictive Coding (DPC) …

Action RecognitionRepresentation LearningSelf-Supervised Action RecognitionSelf-Supervised Action Recognition Linear+2

An Emerging Coding Paradigm VCM: A Scalable Coding Approach Beyond Feature and Signal

2020-01-09 · Sifeng Xia, Kunchangtai Liang, Wenhan Yang, Ling-Yu Duan 외

In this paper, we study a new problem arising from the emerging MPEG standardization effort Video Coding for Machine (VCM), which aims to bridge the gap between visual feature compression and classical video coding. VCM …

Action RecognitionFeature CompressionSSIM