paper-with-me

홈 › Papers

Diverse Video Captioning Through Latent Variable Expansion

2019-10-26 · Huanhou Xiao, Jinglun Shi

Automatically describing video content with text description is challenging but important task, which has been attracting a lot of attention in computer vision community. Previous works mainly strive for the accuracy of the generated sentences, while ignoring the sentences diversity, which is inconsistent with human behavior. In this paper, we aim to caption each video with multiple descriptions and propose a novel framework. Concretely, for a given video, the intermediate latent variables of conventional encode-decode process are utilized as input to the conditional generative adversarial network (CGAN) with the purpose of generating diverse sentences. We adopt different Convolutional Neural Networks (CNNs) as our generator that produces descriptions conditioned on latent variables and discriminator that assesses the quality of generated sentences. Simultaneously, a novel DCE metric is designed to assess the diverse captions. We evaluate our method on the benchmark datasets, where it demonstrates its ability to generate diverse descriptions and achieves superior results against other state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:1910.12019

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityGenerative Adversarial NetworkVideo Captioning

Similar Papers 제목 키워드 기반

From Deterministic to Generative: Multi-Modal Stochastic RNNs for Video Captioning

2017-08-08 · Jingkuan Song, Yuyu Guo, Lianli Gao, Xuelong. Li 외

Video captioning in essential is a complex natural process, which is affected by various uncertainties stemming from video content, subjective judgment, etc. In this paper we build on the recent progress in using encoder…

DecoderVideo Captioning

Diverse Image Captioning with Context-Object Split Latent Spaces

2020-11-02 · NeurIPS 2020 12 · Shweta Mahajan, Stefan Roth

Diverse image captioning models aim to learn one-to-many mappings that are innate to cross-domain datasets, such as of images and texts. Current methods for this task are based on generative latent variable models, e.g. …

DiversityImage CaptioningObject

Variational Structured Semantic Inference for Diverse Image Captioning

2019-12-01 · NeurIPS 2019 12 · Fuhai Chen, Rongrong Ji, Jiayi Ji, Xiaoshuai Sun 외

Despite the exciting progress in image captioning, generating diverse captions for a given image remains as an open problem. Existing methods typically apply generative models such as Variational Auto-Encoder to diversif…

DecoderDiversityImage Captioning

Exact Adversarial Attack to Image Captioning via Structured Output Learning with Latent Variables

2019-05-10 · CVPR 2019 6 · Yan Xu, Baoyuan Wu, Fumin Shen, Yanbo Fan 외

In this work, we study the robustness of a CNN+RNN based image captioning system being subjected to adversarial noises. We propose to fool an image captioning system to generate some targeted partial captions for an imag…

Adversarial AttackImage Captioning

Weakly Supervised Dense Video Captioning

2017-04-05 · CVPR 2017 7 · Zhiqiang Shen, Jianguo Li, Zhou Su, Minjun Li 외

This paper focuses on a novel and challenging vision task, dense video captioning, which aims to automatically describe a video clip with multiple informative and diverse caption sentences. The proposed method is trained…

Dense Video CaptioningLanguage ModelingLanguage ModellingMulti-Label Learning+2