Deep Learning for Video Classification and Captioning
Accelerated by the tremendous increase in Internet bandwidth and storage space, video data has been generated, published and spread explosively, becoming an indispensable part of today's big data. In this paper, we focus on reviewing two lines of research aiming to stimulate the comprehension of videos with deep learning: video classification and video captioning. While video classification concentrates on automatically labeling video clips based on their semantic contents like human actions or complex events, video captioning attempts to generate a complete and natural sentence, enriching the single label as in video classification, to capture the most informative dynamics in videos. In addition, we also provide a review of popular benchmarks and competitions, which are critical for evaluating the technical progress of this vibrant field.
Code (1)
Tasks
ClassificationDeep LearningGeneral ClassificationSentenceVideo CaptioningVideo ClassificationSimilar Papers 제목 키워드 기반
PIC 4th Challenge: Semantic-Assisted Multi-Feature Encoding and Multi-Head Decoding for Dense Video Captioning
The task of Dense Video Captioning (DVC) aims to generate captions with timestamps for multiple events in one video. Semantic information plays an important role for both localization and description of DVC. We present a…
Dense Video CaptioningVideo CaptioningTechnical Report for Soccernet 2023 -- Dense Video Captioning
In the task of dense video captioning of Soccernet dataset, we propose to generate a video caption of each soccer action and locate the timestamp of the caption. Firstly, we apply Blip as our video caption framework to g…
Dense Video CaptioningVideo CaptioningSpatio-Temporal Attention Models for Grounded Video Captioning
Automatic video captioning is challenging due to the complex interactions in dynamic real scenes. A comprehensive system would ultimately localize and track the objects, actions and interactions present in a video and ge…
image-classificationImage ClassificationTemporal LocalizationVideo CaptioningHuman Action Sequence Classification
This paper classifies human action sequences from videos using a machine translation model. In contrast to classical human action classification which outputs a set of actions, our method output a sequence of action in t…
Action ClassificationAction LocalizationAction SegmentationClassification+4ReGen: A good Generative Zero-Shot Video Classifier Should be Rewarded
This paper sets out to solve the following problem: How can we turn a generative video captioning model into an open-world video/action classification model? Video captioning models can naturally produce open-ended f…
Action ClassificationAction RecognitionDescriptivereinforcement-learning+2