paper-with-me

홈 › Papers

Image Captioning with Semantic Attention

2016-03-12 · CVPR 2016 6 · Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, Jiebo Luo

Automatically generating a natural language description of an image has attracted interests recently both because of its importance in practical applications and because it connects two major artificial intelligence fields: computer vision and natural language processing. Existing approaches are either top-down, which start from a gist of an image and convert it into words, or bottom-up, which come up with words describing various aspects of an image and then combine them. In this paper, we propose a new algorithm that combines both approaches through a model of semantic attention. Our algorithm learns to selectively attend to semantic concept proposals and fuse them into hidden states and outputs of recurrent neural networks. The selection and fusion form a feedback connecting the top-down and bottom-up computation. We evaluate our algorithm on two public benchmarks: Microsoft COCO and Flickr30K. Experimental results show that our algorithm significantly outperforms the state-of-the-art approaches consistently across different evaluation metrics.

📄 PDF Abstract BibTeX arXiv:1603.03925

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

Entangled Transformer for Image Captioning

2019-10-01 · ICCV 2019 10 · Guang Li, Linchao Zhu, Ping Liu, Yi Yang

In image captioning, the typical attention mechanisms are arduous to identify the equivalent visual signals especially when predicting highly abstract words. This phenomenon is known as the semantic gap between vision an…

Image Captioning

Seeing with Humans: Gaze-Assisted Neural Image Captioning

2016-08-18 · Yusuke Sugano, Andreas Bulling

Gaze reflects how humans process visual scenes and is therefore increasingly used in computer vision systems. Previous works demonstrated the potential of gaze for object-centric tasks, such as object localization and re…

Image CaptioningObjectObject LocalizationScene Recognition+1

Stacked Cross-modal Feature Consolidation Attention Networks for Image Captioning

2023-02-08 · Mozhgan PourKeshavarz, Shahabedin Nabavi, Mohsen Ebrahimi Moghaddam, Mehrnoush Shamsfard

Recently, the attention-enriched encoder-decoder framework has aroused great interest in image captioning due to its overwhelming progress. Many visual attention models directly leverage meaningful regions to generate im…

Caption GenerationDecoderImage Captioning

Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation

2025-06-03 · Israa A. Albadarneh, Bassam H. Hammo, Omar S. Al-Kadi

Image captioning involves generating textual descriptions from input images, bridging the gap between computer vision and natural language processing. Recent advancements in transformer-based models have significantly im…

Caption GenerationImage CaptioningScene UnderstandingSurvey

Dual Attention on Pyramid Feature Maps for Image Captioning

2020-11-02 · Litao Yu, Jian Zhang, Qiang Wu

Generating natural sentences from images is a fundamental learning task for visual-semantic understanding in multimedia. In this paper, we propose to apply dual attention on pyramid image feature maps to fully explore th…

DescriptiveImage Captioning