paper-with-me

Papers

Video Question Answering with Iterative Video-Text Co-Tokenization

2022-08-01 · AJ Piergiovanni, Kairo Morton, Weicheng Kuo, Michael S. Ryoo, Anelia Angelova

Video question answering is a challenging task that requires understanding jointly the language input, the visual information in individual video frames, as well as the temporal information about the events occurring in the video. In this paper, we propose a novel multi-stream video encoder for video question answering that uses multiple video inputs and a new video-text iterative co-tokenization approach to answer a variety of questions related to videos. We experimentally evaluate the model on several datasets, such as MSRVTT-QA, MSVD-QA, IVQA, outperforming the previous state-of-the-art by large margins. Simultaneously, our model reduces the required GFLOPs from 150-360 to only 67, producing a highly efficient video question answering model.

📄 PDF Abstract BibTeX arXiv:2208.00934

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVideo Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

NEWSKVQA: Knowledge-Aware News Video Question Answering

2022-02-08 · Pranay Gupta, Manish Gupta

Answering questions in the context of videos can be helpful in video indexing, video retrieval systems, video summarization, learning management systems and surveillance video analysis. Although there exists a large body…

Common Sense ReasoningManagementMultiple-choiceQuestion Answering+6

Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering

2019-04-08 · CVPR 2019 6 · Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang 외

In this paper, we propose a novel end-to-end trainable Video Question Answering (VideoQA) framework with three major components: 1) a new heterogeneous memory which can effectively learn global context information from a…

Question AnsweringVideo Question AnsweringVisual Question Answering (VQA)

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

2025-07-01 · Anurag Arnab, Ahmet Iscen, Mathilde Caron, Alireza Fathi 외 arxiv

Despite recent advances in Vision-Language Models (VLMs), long-video understanding remains a challenging problem. Although state-of-the-art long-context VLMs can process around 1000 input frames, they still struggle to e…

Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos

2025-12-19 · Henghui Du, Chunjie Zhang, Xi Chen, Chang Zhou 외 arxiv

Long Video Question-Answering (LVQA) presents a significant challenge for Multi-modal Large Language Models (MLLMs) due to immense context and overloaded information, which could also lead to prohibitive memory consumpti…

Neural Reasoning, Fast and Slow, for Video Question Answering

2019-07-10 · Thao Minh Le, Vuong Le, Svetha Venkatesh, Truyen Tran

What does it take to design a machine that learns to answer natural questions about a video? A Video QA system must simultaneously understand language, represent visual content over space-time, and iteratively transform …

Natural QuestionsQuestion AnsweringVideo Question AnsweringVisual Question Answering+1