paper-with-me

Papers

StyleNet: Generating Attractive Visual Captions With Styles

2017-07-01 · CVPR 2017 7 · Chuang Gan, Zhe Gan, Xiaodong He, Jianfeng Gao, Li Deng

We propose a novel framework named StyleNet to address the task of generating attractive captions for images and videos with different styles. To this end, we devise a novel model component, named factored LSTM, which automatically distills the style factors in the monolingual text corpus. Then at runtime, we can explicitly control the style in the caption generation process so as to produce attractive visual captions with the desired style. Our approach achieves this goal by leveraging two sets of data: 1) factual image/video-caption paired data, and 2) stylized monolingual text data (e.g., romantic and humorous sentences). We show experimentally that StyleNet outperforms existing approaches for generating visual captions with different styles, measured in both automatic and human evaluation metrics on the newly collected FlickrStyle10K image caption dataset, which contains 10K Flickr images with corresponding humorous and romantic captions.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Caption Generation

Similar Papers 제목 키워드 기반

SemStyle: Learning to Generate Stylised Image Captions using Unaligned Text

2018-05-18 · CVPR 2018 6 · Alexander Mathews, Lexing Xie, Xuming He

Linguistic style is an essential part of written communication, with the power to affect both clarity and attractiveness. With recent advances in vision and language, we can start to tackle the problem of generating imag…

DescriptiveImage CaptioningLanguage ModelingLanguage Modelling

Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences

2023-07-31 · Dingyi Yang, Hongyu Chen, Xinglin Hou, Tiezheng Ge 외

Stylized visual captioning aims to generate image or video descriptions with specific styles, making them more attractive and emotionally appropriate. One major challenge with this task is the lack of paired stylized cap…

DecoderImage CaptioningLanguage Modelling

Image Captioning at Will: A Versatile Scheme for Effectively Injecting Sentiments into Image Descriptions

2018-01-30 · Quanzeng You, Hailin Jin, Jiebo Luo

Automatic image captioning has recently approached human-level performance due to the latest advances in computer vision and natural language understanding. However, most of the current models can only generate plain fac…

Image CaptioningNatural Language Understanding

ADS-Cap: A Framework for Accurate and Diverse Stylized Captioning with Unpaired Stylistic Corpora

2023-08-02 · Kanzhi Cheng, Zheng Ma, Shi Zong, Jianbing Zhang 외

Generating visually grounded image captions with specific linguistic styles using unpaired stylistic corpora is a challenging task, especially since we expect stylized captions with a wide variety of stylistic patterns. …

Contrastive LearningDiversityImage Captioning

CapOnImage: Context-driven Dense-Captioning on Image

2022-04-27 · Yiqi Gao, Xinglin Hou, Yuanmeng Zhang, Tiezheng Ge 외

Existing image captioning systems are dedicated to generating narrative captions for images, which are spatially detached from the image in presentation. However, texts can also be used as decorations on the image to hig…

Dense CaptioningDiversityImage Captioning