paper-with-me

홈 › Papers

Coherent Multi-Sentence Video Description with Variable Level of Detail

2014-03-24 · Anna Senina, Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Sikandar Amin, Mykhaylo Andriluka, Manfred Pinkal, Bernt Schiele

Humans can easily describe what they see in a coherent way and at varying level of detail. However, existing approaches for automatic video description are mainly focused on single sentence generation and produce descriptions at a fixed level of detail. In this paper, we address both of these limitations: for a variable level of detail we produce coherent multi-sentence descriptions of complex videos. We follow a two-step approach where we first learn to predict a semantic representation (SR) from video and then generate natural language descriptions from the SR. To produce consistent multi-sentence descriptions, we model across-sentence consistency at the level of the SR by enforcing a consistent topic. We also contribute both to the visual recognition of objects proposing a hand-centric approach as well as to the robust generation of sentences using a word lattice. Human judges rate our multi-sentence descriptions as more readable, correct, and relevant than related work. To understand the difference between more detailed and shorter descriptions, we collect and analyze a video description corpus of three levels of detail.

📄 PDF Abstract BibTeX arXiv:1403.6173

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceVideo Description

Similar Papers 제목 키워드 기반

Move Forward and Tell: A Progressive Generator of Video Descriptions

2018-07-26 · ECCV 2018 9 · Yilei Xiong, Bo Dai, Dahua Lin

We present an efficient framework that can generate a coherent paragraph to describe a given video. Previous works on video captioning usually focus on video clips. They typically treat an entire video as a whole and gen…

DescriptiveSentenceVideo Captioning

Adversarial Inference for Multi-Sentence Video Description

2018-12-13 · CVPR 2019 6 · Jae Sung Park, Marcus Rohrbach, Trevor Darrell, Anna Rohrbach

While significant progress has been made in the image captioning task, video description is still in its infancy due to the complex nature of video data. Generating multi-sentence descriptions for long videos is even mor…

DiversityImage CaptioningSentenceVideo Description

Identity-Aware Multi-Sentence Video Description

2020-08-22 · ECCV 2020 8 · Jae Sung Park, Trevor Darrell, Anna Rohrbach

Standard video and movie description tasks abstract away from person identities, thus failing to link identities across sentences. We propose a multi-sentence Identity-Aware Video Description task, which overcomes this l…

Gender PredictionSentenceVideo Description

Diverse Video Captioning Through Latent Variable Expansion

2019-10-26 · Huanhou Xiao, Jinglun Shi

Automatically describing video content with text description is challenging but important task, which has been attracting a lot of attention in computer vision community. Previous works mainly strive for the accuracy of …

DiversityGenerative Adversarial NetworkVideo Captioning

MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph Captioning

2020-05-11 · ACL 2020 6 · Jie Lei, Li-Wei Wang, Yelong Shen, Dong Yu 외

Generating multi-sentence descriptions for videos is one of the most challenging captioning tasks due to its high requirements for not only visual relevance but also discourse-based coherence across the sentences in the …

SentenceVideo Captioning