paper-with-me

홈 › Papers

Adversarial Multimodal Network for Movie Question Answering

2019-06-24 · Zhaoquan Yuan, Siyuan Sun, Lixin Duan, Xiao Wu, Changsheng Xu

Visual question answering by using information from multiple modalities has attracted more and more attention in recent years. However, it is a very challenging task, as the visual content and natural language have quite different statistical properties. In this work, we present a method called Adversarial Multimodal Network (AMN) to better understand video stories for question answering. In AMN, as inspired by generative adversarial networks, we propose to learn multimodal feature representations by finding a more coherent subspace for video clips and the corresponding texts (e.g., subtitles and questions). Moreover, we introduce a self-attention mechanism to enforce the so-called consistency constraints in order to preserve the self-correlation of visual cues of the original video clips in the learned multimodal representations. Extensive experiments on the MovieQA dataset show the effectiveness of our proposed AMN over other published state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:1906.09844

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVideo Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark

2026-01-05 · Shaden Shaar, Bradon Thymes, Sirawut Chaixanien, Claire Cardie 외 arxiv

Understanding real-world videos such as movies requires integrating visual and dialogue cues. Yet existing VideoQA benchmarks struggle to capture this multimodal reasoning and, given the difficulty of evaluating free-for…

Multimodal ReasoningVisual Reasoning

MoVQA: A Benchmark of Versatile Question-Answering for Long-Form Movie Understanding

2023-12-08 · Hongjie Zhang, Yi Liu, Lu Dong, Yifei HUANG 외

While several long-form VideoQA datasets have been introduced, the length of both videos used to curate questions and sub-clips of clues leveraged to answer those questions have not yet reached the criteria for genuine l…

FormQuestion AnsweringVideo Question AnsweringVideo Understanding

Multimodal Dual Attention Memory for Video Story Question Answering

2018-09-21 · ECCV 2018 9 · Kyung-Min Kim, Seong-Ho Choi, Jin-Hwa Kim, Byoung-Tak Zhang

We propose a video story question-answering (QA) architecture, Multimodal Dual Attention Memory (MDAM). The key idea is to use a dual attention mechanism with late fusion. MDAM uses self-attention to learn the latent con…

Question Answering

Learning Video Context as Interleaved Multimodal Sequences

2024-07-31 · Kevin Qinghong Lin, Pengchuan Zhang, Difei Gao, Xide Xia 외

Narrative videos, such as movies, pose significant challenges in video understanding due to their rich contexts (characters, dialogues, storylines) and diverse demands (identify who, relationship, and reason). In this pa…

Language ModelingLanguage ModellingQuestion AnsweringText Retrieval+5

Progressive Attention Memory Network for Movie Story Question Answering

2019-04-18 · CVPR 2019 6 · Junyeong Kim, Minuk Ma, Kyung-Su Kim, Sungjin Kim 외

This paper proposes the progressive attention memory network (PAMN) for movie story question answering (QA). Movie story QA is challenging compared to VQA in two aspects: (1) pinpointing the temporal parts relevant to an…

Question AnsweringVideo Story QAVisual Question Answering (VQA)