paper-with-me

홈 › Papers

ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering

2019-06-06 · Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, DaCheng Tao

Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to the video domain for video question answering (VideoQA). Compared to the image domain where large scale and fully annotated benchmark datasets exists, VideoQA datasets are limited to small scale and are automatically generated, etc. These limitations restrict their applicability in practice. Here we introduce ActivityNet-QA, a fully annotated and large scale VideoQA dataset. The dataset consists of 58,000 QA pairs on 5,800 complex web videos derived from the popular ActivityNet dataset. We present a statistical analysis of our ActivityNet-QA dataset and conduct extensive experiments on it by comparing existing VideoQA baselines. Moreover, we explore various video representation strategies to improve VideoQA performance, especially for long videos. The dataset is available at https://github.com/MILVLG/activitynet-qa

📄 PDF Abstract BibTeX arXiv:1906.02467

Code (1)

MILVLG/activitynet-qa 공식 구현

Tasks

Question AnsweringVideo Question AnsweringVisual Question Answering (VQA)Zero-Shot Video Question Answer

Similar Papers 제목 키워드 기반

ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding

2015-06-01 · CVPR 2015 6 · Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, Juan Carlos Niebles

In spite of many dataset efforts for human action recognition, current computer vision algorithms are still severely limited in terms of the variability and complexity of the actions that they can recognize. This is in p…

Action DetectionAction RecognitionActivity DetectionGeneral Classification+2

VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs

2024-09-30 · Ruotong Liao, Max Erler, Huiyu Wang, Guangyao Zhai 외

In the video-language domain, recent works in leveraging zero-shot Large Language Model-based reasoning for video understanding have become competitive challengers to previous end-to-end models. However, long video under…

EgoSchemaLanguage ModellingLarge Language ModelQuestion Answering+3

The ActivityNet Large-Scale Activity Recognition Challenge 2018 Summary

2018-08-11 · Bernard Ghanem, Juan Carlos Niebles, Cees Snoek, Fabian Caba Heilbron 외

The 3rd annual installment of the ActivityNet Large- Scale Activity Recognition Challenge, held as a full-day workshop in CVPR 2018, focused on the recognition of daily life, high-level, goal-oriented activities from use…

Activity Recognition

Locate before Answering: Answer Guided Question Localization for Video Question Answering

2022-10-05 · Tianwen Qian, Ran Cui, Jingjing Chen, Pai Peng 외

Video question answering (VideoQA) is an essential task in vision-language understanding, which has attracted numerous research attention recently. Nevertheless, existing works mostly achieve promising performances on sh…

Question AnsweringVideo Question Answering

Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model

2024-06-15 · Lu Xu, Sijie Zhu, Chunyuan Li, Chia-Wen Kuo 외

The emerging video LMMs (Large Multimodal Models) have achieved significant improvements on generic video understanding in the form of VQA (Visual Question Answering), where the raw videos are captured by cameras. Howeve…

Question AnsweringVideo UnderstandingVisual Question AnsweringVisual Question Answering (VQA)