A Review of Deep Learning for Video Captioning
Video captioning (VC) is a fast-moving, cross-disciplinary area of research that bridges work in the fields of computer vision, natural language processing (NLP), linguistics, and human-computer interaction. In essence, VC involves understanding a video and describing it with language. Captioning is used in a host of applications from creating more accessible interfaces (e.g., low-vision navigation) to video question answering (V-QA), video retrieval and content generation. This survey covers deep learning-based VC, including but, not limited to, attention-based architectures, graph networks, reinforcement learning, adversarial networks, dense video captioning (DVC), and more. We discuss the datasets and evaluation metrics used in the field, and limitations, applications, challenges, and future directions for VC.
Code (0)
등록된 구현이 없습니다.
Tasks
Deep LearningDense Video CaptioningQuestion AnsweringRetrievalVideo CaptioningVideo Question AnsweringVideo RetrievalSimilar Papers 제목 키워드 기반
Deep Learning for Video Classification and Captioning
Accelerated by the tremendous increase in Internet bandwidth and storage space, video data has been generated, published and spread explosively, becoming an indispensable part of today's big data. In this paper, we focus…
ClassificationDeep LearningGeneral ClassificationSentence+2Video Captioning: a comparative review of where we are and which could be the route
Video captioning is the process of describing the content of a sequence of images capturing its semantic relationships and meanings. Dealing with this task with a single image is arduous, not to mention how difficult it …
Video CaptioningRecent Advances in Video Question Answering: A Review of Datasets and Methods
Video Question Answering (VQA) is a recent emerging challenging task in the field of Computer Vision. Several visual information retrieval techniques like Video Captioning/Description and Video-guided Machine Translation…
Information RetrievalMachine TranslationQuestion AnsweringRetrieval+6Dense Video Captioning: A Survey of Techniques, Datasets and Evaluation Protocols
Untrimmed videos have interrelated events, dependencies, context, overlapping events, object-object interactions, domain specificity, and other semantics that are worth highlighting while describing a video in natural la…
Caption GenerationDense Video CaptioningDiversitySentence+2Vision-Language Pre-training: Basics, Recent Advances, and Future Trends
This paper surveys vision-language pre-training (VLP) methods for multimodal intelligence that have been developed in the last few years. We group these approaches into three categories: ($i$) VLP for image-text tasks, s…
Few-Shot LearningImage Captioningimage-classificationImage Classification+12