ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to the video domain for video question answering (VideoQA). Compared to the image domain where large scale and fully annotated benchmark datasets exists, VideoQA datasets are limited to small scale and are automatically generated, etc. These limitations restrict their applicability in practice. Here we introduce ActivityNet-QA, a fully annotated and large scale VideoQA dataset. The dataset consists of 58,000 QA pairs on 5,800 complex web videos derived from the popular ActivityNet dataset. We present a statistical analysis of our ActivityNet-QA dataset and conduct extensive experiments on it by comparing existing VideoQA baselines. Moreover, we explore various video representation strategies to improve VideoQA performance, especially for long videos. The dataset is available at https://github.com/MILVLG/activitynet-qa
Code (1)
Tasks
Question AnsweringVideo Question AnsweringVisual Question Answering (VQA)Zero-Shot Video Question AnswerSimilar Papers 제목 키워드 기반
ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding
In spite of many dataset efforts for human action recognition, current computer vision algorithms are still severely limited in terms of the variability and complexity of the actions that they can recognize. This is in p…
Action DetectionAction RecognitionActivity DetectionGeneral Classification+2VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs
In the video-language domain, recent works in leveraging zero-shot Large Language Model-based reasoning for video understanding have become competitive challengers to previous end-to-end models. However, long video under…
EgoSchemaLanguage ModellingLarge Language ModelQuestion Answering+3The ActivityNet Large-Scale Activity Recognition Challenge 2018 Summary
The 3rd annual installment of the ActivityNet Large- Scale Activity Recognition Challenge, held as a full-day workshop in CVPR 2018, focused on the recognition of daily life, high-level, goal-oriented activities from use…
Activity RecognitionLocate before Answering: Answer Guided Question Localization for Video Question Answering
Video question answering (VideoQA) is an essential task in vision-language understanding, which has attracted numerous research attention recently. Nevertheless, existing works mostly achieve promising performances on sh…
Question AnsweringVideo Question AnsweringBeyond Raw Videos: Understanding Edited Videos with Large Multimodal Model
The emerging video LMMs (Large Multimodal Models) have achieved significant improvements on generic video understanding in the form of VQA (Visual Question Answering), where the raw videos are captured by cameras. Howeve…
Question AnsweringVideo UnderstandingVisual Question AnsweringVisual Question Answering (VQA)