paper-with-me

Papers

MovieQA: Understanding Stories in Movies through Question-Answering

2015-12-09 · CVPR 2016 6 · Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, Sanja Fidler

We introduce the MovieQA dataset which aims to evaluate automatic story comprehension from both video and text. The dataset consists of 14,944 questions about 408 movies with high semantic diversity. The questions range from simpler "Who" did "What" to "Whom", to "Why" and "How" certain events occurred. Each question comes with a set of five possible answers; a correct one and four deceiving answers provided by human annotators. Our dataset is unique in that it contains multiple sources of information -- video clips, plots, subtitles, scripts, and DVS. We analyze our data through various statistics and methods. We further extend existing QA techniques to show that question-answering with such open-ended semantics is hard. We make this data set public along with an evaluation benchmark to encourage inspiring work in this challenging domain.

📄 PDF Abstract BibTeX arXiv:1512.02902

Code (1)

makarandtapaswi/MovieQA_benchmark

Tasks

DiversityQuestion Answering

Similar Papers 제목 키워드 기반

Movie Question Answering: Remembering the Textual Cues for Layered Visual Contents

2018-04-25 · Bo Wang, Youjiang Xu, Yahong Han, Richang Hong

Movies provide us with a mass of visual content as well as attracting stories. Existing methods have illustrated that understanding movie stories through only visual content is still a hard problem. In this paper, for an…

Question AnsweringVideo Question Answering

Learning Video Context as Interleaved Multimodal Sequences

2024-07-31 · Kevin Qinghong Lin, Pengchuan Zhang, Difei Gao, Xide Xia 외

Narrative videos, such as movies, pose significant challenges in video understanding due to their rich contexts (characters, dialogues, storylines) and diverse demands (identify who, relationship, and reason). In this pa…

Language ModelingLanguage ModellingQuestion AnsweringText Retrieval+5

StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification

2024-11-11 · Yichen He, Yuan Lin, Jianchao Wu, Hanchong Zhang 외

Existing large vision-language models (LVLMs) are largely limited to processing short, seconds-long videos and struggle with generating coherent descriptions for extended video spanning minutes or more. Long video descri…

Large Language ModelMultimodal Large Language ModelMultiple-choiceVideo Description

Neural Event Extraction from Movies Description

2018-06-01 · WS 2018 6 · Alex Tozzo, Dejan Jovanovi{\'c}, Mohamed Amer

We present a novel approach for event extraction and abstraction from movie descriptions. Our event frame consists of {``}who{''}, {``}did what{''} {``}to whom{''}, {``}where{''}, and {``}when{''}. We formulate our probl…

Event ExtractionMachine TranslationQuestion AnsweringStory Completion+1

Speaker Naming in Movies

2018-09-24 · NAACL 2018 6 · Mahmoud Azab, Mingzhe Wang, Max Smith, Noriyuki Kojima 외

We propose a new model for speaker naming in movies that leverages visual, textual, and acoustic modalities in an unified optimization framework. To evaluate the performance of our model, we introduce a new dataset consi…