paper-with-me

홈 › Papers

BDIQA: A New Dataset for Video Question Answering to Explore Cognitive Reasoning through Theory of Mind

2024-02-12 · Yuanyuan Mao, Xin Lin, Qin Ni, Liang He

As a foundational component of cognitive intelligence, theory of mind (ToM) can make AI more closely resemble human thought processes, thereby enhancing their interaction and collaboration with human. In particular, it can significantly improve a model's comprehension of videos in complex scenes. However, current video question answer (VideoQA) datasets focus on studying causal reasoning within events few of them genuinely incorporating human ToM. Consequently, there is a lack of development in ToM reasoning tasks within the area of VideoQA. This paper presents BDIQA, the first benchmark to explore the cognitive reasoning capabilities of VideoQA models in the context of ToM. BDIQA is inspired by the cognitive development of children's ToM and addresses the current deficiencies in machine ToM within datasets and tasks. Specifically, it offers tasks at two difficulty levels, assessing Belief, Desire and Intention (BDI) reasoning in both simple and complex scenarios. We conduct evaluations on several mainstream methods of VideoQA and diagnose their capabilities with zero shot, few shot and supervised learning. We find that the performance of pre-trained models on cognitive reasoning tasks remains unsatisfactory. To counter this challenge, we undertake thorough analysis and experimentation, ultimately presenting two guidelines to enhance cognitive reasoning derived from ablation analysis.

📄 PDF Abstract BibTeX arXiv:2402.07402

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVideo Question Answering

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

NEWSKVQA: Knowledge-Aware News Video Question Answering

2022-02-08 · Pranay Gupta, Manish Gupta

Answering questions in the context of videos can be helpful in video indexing, video retrieval systems, video summarization, learning management systems and surveillance video analysis. Although there exists a large body…

Common Sense ReasoningManagementMultiple-choiceQuestion Answering+6

The Forgettable-Watcher Model for Video Question Answering

2017-05-03 · Hongyang Xue, Zhou Zhao, Deng Cai

A number of visual question answering approaches have been proposed recently, aiming at understanding the visual scenes by answering the natural language questions. While the image question answering has drawn significan…

modelQuestion AnsweringQuestion GenerationQuestion-Generation+3

TutorialVQA: Question Answering Dataset for Tutorial Videos

2019-12-02 · LREC 2020 5 · Anthony Colas, Seokhwan Kim, Franck Dernoncourt, Siddhesh Gupte 외

Despite the number of currently available datasets on video question answering, there still remains a need for a dataset involving multi-step and non-factoid answers. Moreover, relying on video transcripts remains an und…

Question AnsweringVideo Question Answering

ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering

2019-06-06 · Zhou Yu, Dejing Xu, Jun Yu, Ting Yu 외

Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to the video domain for video question answ…

Question AnsweringVideo Question AnsweringVisual Question Answering (VQA)Zero-Shot Video Question Answer

Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports

2024-01-03 · Haopeng Li, Andong Deng, Jun Liu, Hossein Rahmani 외

Reasoning over sports videos for question answering is an important task with numerous applications, such as player training and information retrieval. However, this task has not been explored due to the lack of relevant…

Action UnderstandingcounterfactualInformation RetrievalQuestion Answering+2