paper-with-me

홈 › Papers

Multimodal Dual Attention Memory for Video Story Question Answering

2018-09-21 · ECCV 2018 9 · Kyung-Min Kim, Seong-Ho Choi, Jin-Hwa Kim, Byoung-Tak Zhang

We propose a video story question-answering (QA) architecture, Multimodal Dual Attention Memory (MDAM). The key idea is to use a dual attention mechanism with late fusion. MDAM uses self-attention to learn the latent concepts in scene frames and captions. Given a question, MDAM uses the second attention over these latent concepts. Multimodal fusion is performed after the dual attention processes (late fusion). Using this processing pipeline, MDAM learns to infer a high-level vision-language joint representation from an abstraction of the full video content. We evaluate MDAM on PororoQA and MovieQA datasets which have large-scale QA annotations on cartoon videos and movies, respectively. For both datasets, MDAM achieves new state-of-the-art results with significant margins compared to the runner-up models. We confirm the best performance of the dual attention mechanism combined with late fusion by ablation studies. We also perform qualitative analysis by visualizing the inference mechanisms of MDAM.

📄 PDF Abstract BibTeX arXiv:1809.07999

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Synopses of Movie Narratives: a Video-Language Dataset for Story Understanding

2022-03-11 · Yidan Sun, Qin Chao, Yangfeng Ji, Boyang Li

Despite recent advances of AI, story understanding remains an open and under-investigated problem. We collect, preprocess, and publicly release a video-language story dataset, Synopses of Movie Narratives (SyMoN), contai…

RetrievalText RetrievalVideo-Text Retrieval

MMT: Image-guided Story Ending Generation with Multimodal Memory Transformer

2022-10-10 · ACM MM 2022 10 · Dizhan Xue, Shengsheng Qian, Quan Fang, Changsheng Xu

As a specific form of story generation, Image-guided Story Ending Generation (IgSEG) is a recently proposed task of generating a story ending for a given multi-sentence story plot and an ending-related image. Unlike exis…

DecoderImage CaptioningImage-guided Story Ending GenerationSentence+1

DeepStory: Video Story QA by Deep Embedded Memory Networks

2017-07-04 · Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi, Byoung-Tak Zhang

Question-answering (QA) on video contents is a significant challenge for achieving human-level intelligence as it involves both vision and language in real-world settings. Here we demonstrate the possibility of an AI age…

AI AgentQuestion AnsweringTripletVideo Story QA

Hybrid Reasoning Network for Video-based Commonsense Captioning

2021-08-05 · Weijiang Yu, Jian Liang, Lei Ji, Lu Li 외

The task of video-based commonsense captioning aims to generate event-wise captions and meanwhile provide multiple commonsense descriptions (e.g., attribute, effect and intention) about the underlying event in the video.…

AttributeDecoder

Conversational Memory Network for Emotion Recognition in Dyadic Dialogue Videos

2018-06-01 · NAACL 2018 6 · Devamanyu Hazarika, Soujanya Poria, Amir Zadeh, Erik Cambria 외

Emotion recognition in conversations is crucial for the development of empathetic machines. Present methods mostly ignore the role of inter-speaker dependency relations while classifying emotions in conversations. In thi…

Emotion RecognitionEmotion Recognition in Conversation