paper-with-me

Papers

Spatio-Temporal Attention Models for Grounded Video Captioning

2016-10-17 · Mihai Zanfir, Elisabeta Marinoiu, Cristian Sminchisescu

Automatic video captioning is challenging due to the complex interactions in dynamic real scenes. A comprehensive system would ultimately localize and track the objects, actions and interactions present in a video and generate a description that relies on temporal localization in order to ground the visual concepts. However, most existing automatic video captioning systems map from raw video data to high level textual description, bypassing localization and recognition, thus discarding potentially valuable information for content localization and generalization. In this work we present an automatic video captioning model that combines spatio-temporal attention and image classification by means of deep neural network structures based on long short-term memory. The resulting system is demonstrated to produce state-of-the-art results in the standard YouTube captioning benchmark while also offering the advantage of localizing the visual concepts (subjects, verbs, objects), with no grounding supervision, over space and time.

📄 PDF Abstract BibTeX arXiv:1610.04997

Code (0)

등록된 구현이 없습니다.

Tasks

image-classificationImage ClassificationTemporal LocalizationVideo Captioning

Similar Papers 제목 키워드 기반

Spatio-Temporal Ranked-Attention Networks for Video Captioning

2020-01-17 · Anoop Cherian, Jue Wang, Chiori Hori, Tim K. Marks

Generating video descriptions automatically is a challenging task that involves a complex interplay between spatio-temporal visual features and language models. Given that videos consist of spatial (frame-level) features…

Video Captioning

Grounded Video Caption Generation

2024-11-12 · Evangelos Kazakos, Cordelia Schmid, Josef Sivic

We propose a new task, dataset and model for grounded video caption generation. This task unifies captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally consis…

Caption GenerationImage Captioning

Spatio-Temporal Graph for Video Captioning with Knowledge Distillation

2020-03-31 · CVPR 2020 6 · Boxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee 외

Video captioning is a challenging task that requires a deep understanding of visual scenes. State-of-the-art methods generate captions using either scene-level or object-level information but without explicitly modeling …

Knowledge DistillationObjectVideo CaptioningVisual Grounding

Diverse Video Captioning by Adaptive Spatio-temporal Attention

2022-08-19 · Zohreh Ghaderi, Leonard Salewski, Hendrik P. A. Lensch

To generate proper captions for videos, the inference needs to identify relevant concepts and pay attention to the spatial relationships between them as well as to the temporal development in the clip. Our end-to-end enc…

DecoderDiversityText GenerationVideo Captioning

Object-aware Aggregation with Bidirectional Temporal Graph for Video Captioning

2019-06-11 · CVPR 2019 6 · Junchao Zhang, Yuxin Peng

Video captioning aims to automatically generate natural language descriptions of video content, which has drawn a lot of attention recent years. Generating accurate and fine-grained captions needs to not only understand …

ObjectVideo Captioning