paper-with-me

홈 › Papers

DeepStory: Video Story QA by Deep Embedded Memory Networks

2017-07-04 · Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi, Byoung-Tak Zhang

Question-answering (QA) on video contents is a significant challenge for achieving human-level intelligence as it involves both vision and language in real-world settings. Here we demonstrate the possibility of an AI agent performing video story QA by learning from a large amount of cartoon videos. We develop a video-story learning model, i.e. Deep Embedded Memory Networks (DEMN), to reconstruct stories from a joint scene-dialogue video stream using a latent embedding space of observed data. The video stories are stored in a long-term memory component. For a given question, an LSTM-based attention model uses the long-term memory to recall the best question-story-answer triplet by focusing on specific words containing key information. We trained the DEMN on a novel QA dataset of children's cartoon video series, Pororo. The dataset contains 16,066 scene-dialogue pairs of 20.5-hour videos, 27,328 fine-grained sentences for scene description, and 8,913 story-related QA pairs. Our experimental results show that the DEMN outperforms other QA models. This is mainly due to 1) the reconstruction of video stories in a scene-dialogue combined form that utilize the latent embedding and 2) attention. DEMN also achieved state-of-the-art results on the MovieQA benchmark.

📄 PDF Abstract BibTeX arXiv:1707.00836

Code (0)

등록된 구현이 없습니다.

Tasks

AI AgentQuestion AnsweringTripletVideo Story QA

Similar Papers 제목 키워드 기반

StoryMem: Multi-shot Long Video Storytelling with Memory

2025-12-22 · Kaiwen Zhang, Liming Jiang, Angtian Wang, Jacob Zhiyuan Fang 외 arxiv

Visual storytelling requires generating multi-shot videos with cinematic quality and long-range consistency. Inspired by human memory, we propose StoryMem, a paradigm that reformulates long-form video storytelling as ite…

Visual StorytellingStory Generation

A Memory Network Approach for Story-Based Temporal Summarization of 360° Videos

2018-06-01 · CVPR 2018 6 · Sang-ho Lee, Jinyoung Sung, Youngjae Yu, Gunhee Kim

We address the problem of story-based temporal summarization of long 360° videos. We propose a novel memory network model named Past-Future Memory Network (PFMN), in which we first compute the scores of 81 normal field …

Video Summarization

A Memory Network Approach for Story-based Temporal Summarization of 360° Videos

2018-05-08 · CVPR 2018 · Sang-ho Lee, Jinyoung Sung, Youngjae Yu, Gunhee Kim

We address the problem of story-based temporal summarization of long 360{\deg} videos. We propose a novel memory network model named Past-Future Memory Network (PFMN), in which we first compute the scores of 81 normal fi…

Video Summarization

Memory Storyboard: Leveraging Temporal Segmentation for Streaming Self-Supervised Learning from Egocentric Videos

2025-01-21 · Yanlai Yang, Mengye Ren

Self-supervised learning holds the promise to learn good representations from real-world continuous uncurated data streams. However, most existing works in visual self-supervised learning focus on static images or artifi…

Continual LearningContrastive LearningEvent SegmentationSelf-Supervised Learning

TinyHistory: Lightweight Video History Embeddings via Two-Stage Context Learning

2025-12-29 · Lvmin Zhang, Shengqu Cai, Muyang Li, Chong Zeng 외 arxiv

History context is central to autoregressive video generation, driving consistency and storytelling for both commercial models and personal use cases. For example, personal users, offline workflows, and individual-scale …

Video Generation