paper-with-me

홈 › Papers

The Long-Short Story of Movie Description

2015-06-04 · Anna Rohrbach, Marcus Rohrbach, Bernt Schiele

Generating descriptions for videos has many applications including assisting blind people and human-robot interaction. The recent advances in image captioning as well as the release of large-scale movie description datasets such as MPII Movie Description allow to study this task in more depth. Many of the proposed methods for image captioning rely on pre-trained object classifier CNNs and Long-Short Term Memory recurrent networks (LSTMs) for generating descriptions. While image description focuses on objects, we argue that it is important to distinguish verbs, objects, and places in the challenging setting of movie description. In this work we show how to learn robust visual classifiers from the weak annotations of the sentence descriptions. Based on these visual classifiers we learn how to generate a description using an LSTM. We explore different design choices to build and train the LSTM and achieve the best performance to date on the challenging MPII-MD dataset. We compare and analyze our approach and prior work along various dimensions to better understand the key challenges of the movie description task.

📄 PDF Abstract BibTeX arXiv:1506.01698

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningImage DescriptionSentence

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification

2024-11-11 · Yichen He, Yuan Lin, Jianchao Wu, Hanchong Zhang 외

Existing large vision-language models (LVLMs) are largely limited to processing short, seconds-long videos and struggle with generating coherent descriptions for extended video spanning minutes or more. Long video descri…

Large Language ModelMultimodal Large Language ModelMultiple-choiceVideo Description

Captain Cinema: Towards Short Movie Generation

2025-07-24 · Junfei Xiao, Ceyuan Yang, Lvmin Zhang, Shengqu Cai 외 arxiv

We present Captain Cinema, a generation framework for short movie generation. Given a detailed textual description of a movie storyline, our approach firstly generates a sequence of keyframes that outline the entire narr…

MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation

2024-11-22 · CVPR 2025 1 · Weijia Wu, MingYu Liu, Zeyu Zhu, Xi Xia 외

Recent advancements in video generation models, like Stable Video Diffusion, show promising results, but primarily focus on short, single-scene videos. These models struggle with generating long videos that involve multi…

Video Generation

MovieNet: A Holistic Dataset for Movie Understanding

2020-07-21 · ECCV 2020 8 · Qingqiu Huang, Yu Xiong, Anyi Rao, Jiaze Wang 외

Recent years have seen remarkable advances in visual understanding. However, how to understand a story-based long video with artistic styles, e.g. movie, remains challenging. In this paper, we introduce MovieNet -- a hol…

Video Understanding

Synopses of Movie Narratives: a Video-Language Dataset for Story Understanding

2022-03-11 · Yidan Sun, Qin Chao, Yangfeng Ji, Boyang Li

Despite recent advances of AI, story understanding remains an open and under-investigated problem. We collect, preprocess, and publicly release a video-language story dataset, Synopses of Movie Narratives (SyMoN), contai…

RetrievalText RetrievalVideo-Text Retrieval