StyleNet: Generating Attractive Visual Captions With Styles
We propose a novel framework named StyleNet to address the task of generating attractive captions for images and videos with different styles. To this end, we devise a novel model component, named factored LSTM, which automatically distills the style factors in the monolingual text corpus. Then at runtime, we can explicitly control the style in the caption generation process so as to produce attractive visual captions with the desired style. Our approach achieves this goal by leveraging two sets of data: 1) factual image/video-caption paired data, and 2) stylized monolingual text data (e.g., romantic and humorous sentences). We show experimentally that StyleNet outperforms existing approaches for generating visual captions with different styles, measured in both automatic and human evaluation metrics on the newly collected FlickrStyle10K image caption dataset, which contains 10K Flickr images with corresponding humorous and romantic captions.
Code (0)
등록된 구현이 없습니다.
Tasks
Caption GenerationSimilar Papers 제목 키워드 기반
SemStyle: Learning to Generate Stylised Image Captions using Unaligned Text
Linguistic style is an essential part of written communication, with the power to affect both clarity and attractiveness. With recent advances in vision and language, we can start to tackle the problem of generating imag…
DescriptiveImage CaptioningLanguage ModelingLanguage ModellingVisual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences
Stylized visual captioning aims to generate image or video descriptions with specific styles, making them more attractive and emotionally appropriate. One major challenge with this task is the lack of paired stylized cap…
DecoderImage CaptioningLanguage ModellingImage Captioning at Will: A Versatile Scheme for Effectively Injecting Sentiments into Image Descriptions
Automatic image captioning has recently approached human-level performance due to the latest advances in computer vision and natural language understanding. However, most of the current models can only generate plain fac…
Image CaptioningNatural Language UnderstandingADS-Cap: A Framework for Accurate and Diverse Stylized Captioning with Unpaired Stylistic Corpora
Generating visually grounded image captions with specific linguistic styles using unpaired stylistic corpora is a challenging task, especially since we expect stylized captions with a wide variety of stylistic patterns. …
Contrastive LearningDiversityImage CaptioningCapOnImage: Context-driven Dense-Captioning on Image
Existing image captioning systems are dedicated to generating narrative captions for images, which are spatially detached from the image in presentation. However, texts can also be used as decorations on the image to hig…
Dense CaptioningDiversityImage Captioning