3M: Multi-style image caption generation using Multi-modality features under Multi-UPDOWN model
In this paper, we build a multi-style generative model for stylish image captioning which uses multi-modality image features, ResNeXt features and text features generated by DenseCap. We propose the 3M model, a Multi-UPDOWN caption model that encodes multi-modality features and decode them to captions. We demonstrate the effectiveness of our model on generating human-like captions by examining its performance on two datasets, the PERSONALITY-CAPTIONS dataset and the FlickrStyle10K dataset. We compare against a variety of state-of-the-art baselines on various automatic NLP metrics such as BLEU, ROUGE-L, CIDEr, SPICE, etc. A qualitative study has also been done to verify our 3M model can be used for generating different stylized captions.
Code (0)
등록된 구현이 없습니다.
Tasks
Caption GenerationImage CaptioningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
UnMA-CapSumT: Unified and Multi-Head Attention-driven Caption Summarization Transformer
Image captioning is the generation of natural language descriptions of images which have increased immense popularity in the recent past. With this different deep-learning techniques are devised for the development of fa…
Image CaptioningStyleNet: Generating Attractive Visual Captions With Styles
We propose a novel framework named StyleNet to address the task of generating attractive captions for images and videos with different styles. To this end, we devise a novel model component, named factored LSTM, which a…
Caption GenerationMSCap: Multi-Style Image Captioning With Unpaired Stylized Text
In this paper, we propose an adversarial learning network for the task of multi-style image captioning (MSCap) with a standard factual image caption dataset and a multi-stylized language corpus without paired images. How…
Image CaptioningSentenceStyle-Aware Contrastive Learning for Multi-Style Image Captioning
Existing multi-style image captioning methods show promising results in generating a caption with accurate visual content and desired linguistic style. However, existing methods overlook the relationship between linguist…
Contrastive LearningImage CaptioningRetrievalTripletDiverse and Styled Image Captioning Using SVD-Based Mixture of Recurrent Experts
With great advances in vision and natural language processing, the generation of image captions becomes a need. In a recent paper, Mathews, Xie and He [1], extended a new model to generate styled captions by separating s…
Image CaptioningSentence