Memory-augmented Dense Predictive Coding for Video Representation Learning
The objective of this paper is self-supervised learning from video, in particular for representations for action recognition. We make the following contributions: (i) We propose a new architecture and learning framework Memory-augmented Dense Predictive Coding (MemDPC) for the task. It is trained with a predictive attention mechanism over the set of compressed memories, such that any future states can always be constructed by a convex combination of the condense representations, allowing to make multiple hypotheses efficiently. (ii) We investigate visual-only self-supervised video representation learning from RGB frames, or from unsupervised optical flow, or both. (iii) We thoroughly evaluate the quality of learnt representation on four different downstream tasks: action recognition, video retrieval, learning with scarce annotations, and unintentional action classification. In all cases, we demonstrate state-of-the-art or comparable performance over other approaches with orders of magnitude fewer training data.
Code (1)
Tasks
Action ClassificationAction RecognitionOptical Flow EstimationRepresentation LearningRetrievalSelf-Supervised LearningVideo RetrievalSimilar Papers 제목 키워드 기반
Video Representation Learning by Dense Predictive Coding
The objective of this paper is self-supervised learning of spatio-temporal embeddings from video, suitable for human action recognition. We make three contributions: First, we introduce the Dense Predictive Coding (DPC) …
Action RecognitionRepresentation LearningSelf-Supervised Action RecognitionSelf-Supervised Action Recognition Linear+2Semantic and episodic memories in a predictive coding model of the neocortex
Complementary Learning Systems theory holds that intelligent agents need two learning systems. Semantic memory is encoded in the neocortex with dense, overlapping representations and acquires structured knowledge. Episod…
Streaming Dense Video Captioning
An ideal model for dense video captioning -- predicting captions localized temporally in a video -- should be able to handle long input videos, predict rich, detailed textual descriptions, and be able to produce outputs …
Dense Video CaptioningLive Video CaptioningVideo CaptioningDeep Hierarchical Video Compression
Recently, probabilistic predictive coding that directly models the conditional distribution of latent features across successive frames for temporal redundancy removal has yielded promising results. Existing methods usin…
Video CompressionMemory Augmented Multi-Instance Contrastive Predictive Coding for Sequential Recommendation
The sequential recommendation aims to recommend items, such as products, songs and places, to users based on the sequential patterns of their historical records. Most existing sequential recommender models consider the n…
Contrastive LearningSequential Recommendation