An Empirical Study of Language CNN for Image Captioning
Language Models based on recurrent neural networks have dominated recent image caption generation tasks. In this paper, we introduce a Language CNN model which is suitable for statistical language modeling tasks and shows competitive performance in image captioning. In contrast to previous models which predict next word based on one previous word and hidden state, our language CNN is fed with all the previous words and can model the long-range dependencies of history words, which are critical for image captioning. The effectiveness of our approach is validated on two datasets MS COCO and Flickr30K. Our extensive experimental results show that our method outperforms the vanilla recurrent neural network based language models and is competitive with the state-of-the-art methods.
Code (2)
Tasks
Caption GenerationImage CaptioningLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
Leveraging Pre-trained BERT for Audio Captioning
Audio captioning aims at using natural language to describe the content of an audio clip. Existing audio captioning systems are generally based on an encoder-decoder architecture, in which acoustic information is extract…
AudioCapsAudio captioningDecoderLanguage ModellingAttacking Visual Language Grounding with Adversarial Examples: A Case Study on Neural Image Captioning
Visual language grounding is widely studied in modern neural image captioning systems, which typically adopts an encoder-decoder framework consisting of two principal components: a convolutional neural network (CNN) for …
Caption GenerationDecoderImage CaptioningScaling Up Vision-Language Pre-training for Image Captioning
In recent years, we have witnessed significant performance boost in the image captioning task based on vision-language pre-training (VLP). Scale is believed to be an important factor for this advance. However, most exist…
AttributeImage CaptioningImproving Image Captioning with Conditional Generative Adversarial Nets
In this paper, we propose a novel conditional-generative-adversarial-nets-based image captioning framework as an extension of traditional reinforcement-learning (RL)-based encoder-decoder architecture. To deal with the i…
DecoderImage CaptioningReinforcement LearningReinforcement Learning (RL)Are scene graphs good enough to improve Image Captioning?
Many top-performing image captioning models rely solely on object features computed with an object detection model to generate image descriptions. However, recent studies propose to directly use scene graphs to introduce…
DecoderGraph AttentionGraph GenerationImage Captioning+4