paper-with-me

홈 › Papers

WildQA: In-the-Wild Video Question Answering

2022-09-14 · Santiago Castro, Naihao Deng, Pingxuan Huang, Mihai Burzo, Rada Mihalcea

Existing video understanding datasets mostly focus on human interactions, with little attention being paid to the "in the wild" settings, where the videos are recorded outdoors. We propose WILDQA, a video understanding dataset of videos recorded in outside settings. In addition to video question answering (Video QA), we also introduce the new task of identifying visual support for a given question and answer (Video Evidence Selection). Through evaluations using a wide range of baseline models, we show that WILDQA poses new challenges to the vision and language research communities. The dataset is available at https://lit.eecs.umich.edu/wildqa/.

📄 PDF Abstract BibTeX arXiv:2209.06650

Code (0)

등록된 구현이 없습니다.

Tasks

Evidence SelectionQuestion AnsweringVideo Question AnsweringVideo Understanding

Similar Papers 제목 키워드 기반

In-the-Wild Video Question Answering

2022-10-01 · COLING 2022 10 · Santiago Castro, Naihao Deng, Pingxuan Huang, Mihai Burzo 외

Existing video understanding datasets mostly focus on human interactions, with little attention being paid to the “in the wild” settings, where the videos are recorded outdoors. We propose WILDQA, a video understanding d…

Evidence SelectionQuestion AnsweringVideo Question AnsweringVideo Understanding

Vision-Language Pre-training: Basics, Recent Advances, and Future Trends

2022-10-17 · Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang 외

This paper surveys vision-language pre-training (VLP) methods for multimodal intelligence that have been developed in the last few years. We group these approaches into three categories: ($i$) VLP for image-text tasks, s…

Few-Shot LearningImage Captioningimage-classificationImage Classification+12

SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning over Traffic Events

2021-03-29 · CVPR 2021 1 · Li Xu, He Huang, Jun Liu

Traffic event cognition and reasoning in videos is an important task that has a wide range of applications in intelligent transportation, assisted driving, and autonomous vehicles. In this paper, we create a novel datase…

Autonomous VehiclesBenchmarkingCausal InferenceQuestion Answering+2

Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models

2024-06-19 · Zhawnen Chen, Tianchun Wang, Yizhou Wang, Michal Kosinski 외

Can large multimodal models have a human-like ability for emotional and social reasoning, and if so, how does it work? Recent research has discovered emergent theory-of-mind (ToM) reasoning capabilities in large language…

NEWSKVQA: Knowledge-Aware News Video Question Answering

2022-02-08 · Pranay Gupta, Manish Gupta

Answering questions in the context of videos can be helpful in video indexing, video retrieval systems, video summarization, learning management systems and surveillance video analysis. Although there exists a large body…

Common Sense ReasoningManagementMultiple-choiceQuestion Answering+6