paper-with-me

Papers

MoVQA: A Benchmark of Versatile Question-Answering for Long-Form Movie Understanding

2023-12-08 · Hongjie Zhang, Yi Liu, Lu Dong, Yifei HUANG, Zhen-Hua Ling, Yali Wang, LiMin Wang, Yu Qiao

While several long-form VideoQA datasets have been introduced, the length of both videos used to curate questions and sub-clips of clues leveraged to answer those questions have not yet reached the criteria for genuine long-form video understanding. Moreover, their QAs are unduly narrow and modality-biased, lacking a wider view of understanding long-term video content with rich dynamics and complex narratives. To remedy this, we introduce MoVQA, a long-form movie question-answering dataset, and benchmark to assess the diverse cognitive capabilities of multimodal systems rely on multi-level temporal lengths, with considering both video length and clue length. Additionally, to take a step towards human-level understanding in long-form video, versatile and multimodal question-answering is designed from the moviegoer-perspective to assess the model capabilities on various perceptual and cognitive axes.Through analysis involving various baselines reveals a consistent trend: the performance of all methods significantly deteriorate with increasing video and clue length. Meanwhile, our established baseline method has shown some improvements, but there is still ample scope for enhancement on our challenging MoVQA dataset. We expect our MoVQA to provide a new perspective and encourage inspiring works on long-form video understanding research.

📄 PDF Abstract BibTeX arXiv:2312.04817

Code (0)

등록된 구현이 없습니다.

Tasks

FormQuestion AnsweringVideo Question AnsweringVideo Understanding

Similar Papers 제목 키워드 기반

Benchmarks for Pirá 2.0, a Reading Comprehension Dataset about the Ocean, the Brazilian Coast, and Climate Change

2023-09-19 · Paulo Pirozelli, Marcos M. José, Igor Silveira, Flávio Nakasato 외

Pir\'a is a reading comprehension dataset focused on the ocean, the Brazilian coast, and climate change, built from a collection of scientific abstracts and reports on these topics. This dataset represents a versatile la…

Generative Question AnsweringInformation RetrievalMachine Reading ComprehensionMultiple-choice+2

Vamos: Versatile Action Models for Video Understanding

2023-11-22 · Shijie Wang, Qi Zhao, Minh Quan Do, Nakul Agarwal 외

What makes good representations for video understanding, such as anticipating future activities, or answering video-conditioned questions? While earlier approaches focus on end-to-end learning directly from video pixels,…

EgoSchemaHard AttentionLanguage ModellingLarge Language Model+4

VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding

2024-06-18 · Xiang Li, Jian Ding, Mohamed Elhoseiny

We introduce a new benchmark designed to advance the development of general-purpose, large-scale vision-language models for remote sensing images. Although several vision-language datasets in remote sensing have been pro…

Image CaptioningQuestion AnsweringVisual GroundingVisual Question Answering

SEAL: Semantic Attention Learning for Long Video Representation

2024-12-02 · CVPR 2025 1 · Lan Wang, Yujia Chen, Du Tran, Vishnu Naresh Boddeti 외

Long video understanding presents challenges due to the inherent high computational complexity and redundant temporal information. An effective representation for long videos must efficiently process such redundancy whil…

DiversityQuestion AnsweringVideo Question AnsweringVideo Understanding

Long Context Question Answering via Supervised Contrastive Learning

2021-12-16 · NAACL 2022 7 · Avi Caciularu, Ido Dagan, Jacob Goldberger, Arman Cohan

Long-context question answering (QA) tasks require reasoning over a long document or multiple documents. Addressing these tasks often benefits from identifying a set of evidence spans (e.g., sentences), which provide sup…

Contrastive LearningQuestion AnsweringVideo Generation