paper-with-me

Papers

3M: Multi-style image caption generation using Multi-modality features under Multi-UPDOWN model

2021-03-20 · Chengxi Li, Brent Harrison

In this paper, we build a multi-style generative model for stylish image captioning which uses multi-modality image features, ResNeXt features and text features generated by DenseCap. We propose the 3M model, a Multi-UPDOWN caption model that encodes multi-modality features and decode them to captions. We demonstrate the effectiveness of our model on generating human-like captions by examining its performance on two datasets, the PERSONALITY-CAPTIONS dataset and the FlickrStyle10K dataset. We compare against a variety of state-of-the-art baselines on various automatic NLP metrics such as BLEU, ROUGE-L, CIDEr, SPICE, etc. A qualitative study has also been done to verify our 3M model can be used for generating different stylized captions.

📄 PDF Abstract BibTeX arXiv:2103.11186

Code (0)

등록된 구현이 없습니다.

Tasks

Caption GenerationImage Captioning

Methods 이 논문이 사용한 방법론

Average Pooling 설명 없음
Global Average Pooling Global Average Pooling is a pooling operation designed to replace fully connected layers in classical CNNs. The idea is to generate one feature map for each corresponding…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Kaiming Initialization 설명 없음
ResNeXt Block A ResNeXt Block is a type of residual block used as part of the ResNeXt CNN…
Grouped Convolution A Grouped Convolution uses a group of convolutions - multiple kernels per layer - resulting in multiple channel outputs per layer. This leads to wider networks helping a…

Similar Papers 제목 키워드 기반

UnMA-CapSumT: Unified and Multi-Head Attention-driven Caption Summarization Transformer

2024-12-16 · Dhruv Sharma, Chhavi Dhiman, Dinesh Kumar

Image captioning is the generation of natural language descriptions of images which have increased immense popularity in the recent past. With this different deep-learning techniques are devised for the development of fa…

Image Captioning

StyleNet: Generating Attractive Visual Captions With Styles

2017-07-01 · CVPR 2017 7 · Chuang Gan, Zhe Gan, Xiaodong He, Jianfeng Gao 외

We propose a novel framework named StyleNet to address the task of generating attractive captions for images and videos with different styles. To this end, we devise a novel model component, named factored LSTM, which a…

Caption Generation

MSCap: Multi-Style Image Captioning With Unpaired Stylized Text

2019-06-01 · CVPR 2019 6 · Longteng Guo, Jing Liu, Peng Yao, Jiangwei Li 외

In this paper, we propose an adversarial learning network for the task of multi-style image captioning (MSCap) with a standard factual image caption dataset and a multi-stylized language corpus without paired images. How…

Image CaptioningSentence

Style-Aware Contrastive Learning for Multi-Style Image Captioning

2023-01-26 · Yucheng Zhou, Guodong Long

Existing multi-style image captioning methods show promising results in generating a caption with accurate visual content and desired linguistic style. However, existing methods overlook the relationship between linguist…

Contrastive LearningImage CaptioningRetrievalTriplet

Diverse and Styled Image Captioning Using SVD-Based Mixture of Recurrent Experts

2020-07-07 · Marzieh Heidari, Mehdi Ghatee, Ahmad Nickabadi, Arash Pourhasan Nezhad

With great advances in vision and natural language processing, the generation of image captions becomes a need. In a recent paper, Mathews, Xie and He [1], extended a new model to generate styled captions by separating s…

Image CaptioningSentence