paper-with-me

홈 › Papers

Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer Network

2020-12-13 · Jiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen, Gen Luo, Yongjian Wu, Yue Gao, Rongrong Ji

Transformer-based architectures have shown great success in image captioning, where object regions are encoded and then attended into the vectorial representations to guide the caption decoding. However, such vectorial representations only contain region-level information without considering the global information reflecting the entire image, which fails to expand the capability of complex multi-modal reasoning in image captioning. In this paper, we introduce a Global Enhanced Transformer (termed GET) to enable the extraction of a more comprehensive global representation, and then adaptively guide the decoder to generate high-quality captions. In GET, a Global Enhanced Encoder is designed for the embedding of the global feature, and a Global Adaptive Decoder are designed for the guidance of the caption generation. The former models intra- and inter-layer global representation by taking advantage of the proposed Global Enhanced Attention and a layer-wise fusion module. The latter contains a Global Adaptive Controller that can adaptively fuse the global information into the decoder to guide the caption generation. Extensive experiments on MS COCO dataset demonstrate the superiority of our GET over many state-of-the-arts.

📄 PDF Abstract BibTeX arXiv:2012.07061

Code (1)

luo3300612/image-captioning-DLCT pytorch

Tasks

Caption GenerationDecoderImage Captioning

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Multi-Head Attention 설명 없음
Residual Connection 설명 없음
Attention 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

Towards Local Visual Modeling for Image Captioning

2023-02-13 · Yiwei Ma, Jiayi Ji, Xiaoshuai Sun, Yiyi Zhou 외

In this paper, we study the local visual modeling with grid features for image captioning, which is critical for generating accurate and detailed captions. To achieve this target, we propose a Locality-Sensitive Transfor…

Image CaptioningObject Recognition

Multimodal Transformer with Multi-View Visual Representation for Image Captioning

2019-05-20 · Jun Yu, Jing Li, Zhou Yu, Qingming Huang

Image captioning aims to automatically generate a natural language description of a given image, and most state-of-the-art models have adopted an encoder-decoder framework. The framework consists of a convolution neural …

DecoderImage CaptioningMachine TranslationMultimodal Reasoning

DeeCap: Dynamic Early Exiting for Efficient Image Captioning

2022-01-01 · CVPR 2022 1 · Zhengcong Fei, Xu Yan, Shuhui Wang, Qi Tian

Both accuracy and efficiency are crucial for image captioning in real-world scenarios. Although Transformer-based models have gained significant improved captioning performance, their computational cost is very high.…

Image CaptioningImitation Learning

Boosted Attention: Leveraging Human Attention for Image Captioning

2019-03-18 · ECCV 2018 9 · Shi Chen, Qi Zhao

Visual attention has shown usefulness in image captioning, with the goal of enabling a caption model to selectively focus on regions of interest. Existing models typically rely on top-down language information and learn …

Image Captioning

Crisscrossed Captions: Extended Intramodal and Intermodal Semantic Similarity Judgments for MS-COCO

2020-04-30 · EACL 2021 2 · Zarana Parekh, Jason Baldridge, Daniel Cer, Austin Waters 외

By supporting multi-modal retrieval training and evaluation, image captioning datasets have spurred remarkable progress on representation learning. Unfortunately, datasets have limited cross-modal associations: images ar…

Image CaptioningRepresentation LearningRetrievalSemantic Similarity+1