Hierarchical Video Frame Sequence Representation with Deep Convolutional Graph Network
High accuracy video label prediction (classification) models are attributed to large scale data. These data could be frame feature sequences extracted by a pre-trained convolutional-neural-network, which promote the efficiency for creating models. Unsupervised solutions such as feature average pooling, as a simple label-independent parameter-free based method, has limited ability to represent the video. While the supervised methods, like RNN, can greatly improve the recognition accuracy. However, the video length is usually long, and there are hierarchical relationships between frames across events in the video, the performance of RNN based models are decreased. In this paper, we proposes a novel video classification method based on a deep convolutional graph neural network(DCGN). The proposed method utilize the characteristics of the hierarchical structure of the video, and performed multi-level feature extraction on the video frame sequence through the graph network, obtained a video representation re ecting the event semantics hierarchically. We test our model on YouTube-8M Large-Scale Video Understanding dataset, and the result outperforms RNN based benchmarks.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationGraph Neural NetworkVideo ClassificationVideo UnderstandingSimilar Papers 제목 키워드 기반
A Joint Sequence Fusion Model for Video Question Answering and Retrieval
We present an approach named JSFusion (Joint Sequence Fusion) that can measure semantic similarity between any pairs of multimodal sequence data (e.g. a video clip and a language sentence). Our multimodal matching networ…
DecoderMultiple-choiceQuestion AnsweringRetrieval+6Open-Ended Long-Form Video Question Answering via Hierarchical Convolutional Self-Attention Networks
Open-ended video question answering aims to automatically generate the natural-language answer from referenced video contents according to the given question. Currently, most existing approaches focus on short-form video…
Answer GenerationDecoderFormQuestion Answering+1Discriminative Hierarchical Rank Pooling for Activity Recognition
We present hierarchical rank pooling, a video sequence encoding method for activity recognition. It consists of a network of rank pooling functions which captures the dynamics of rich convolutional neural network feature…
Action RecognitionActivity RecognitionTemporal Action LocalizationDiscriminatively Learned Hierarchical Rank Pooling Networks
In this work, we present novel temporal encoding methods for action and activity classification by extending the unsupervised rank pooling temporal encoding method in two ways. First, we present "discriminative rank pool…
Activity RecognitionBilevel OptimizationGeneral ClassificationVisual Sequence Learning in Hierarchical Prediction Networks and Primate Visual Cortex
In this paper we developed a computational hierarchical network model to understand the spatiotemporal sequence learning effects observed in the primate visual cortex. The model is a hierarchical recurrent neural model…