paper-with-me

홈 › Papers

Learning to Guide Decoding for Image Captioning

2018-04-03 · Wenhao Jiang, Lin Ma, Xinpeng Chen, Hanwang Zhang, Wei Liu

Recently, much advance has been made in image captioning, and an encoder-decoder framework has achieved outstanding performance for this task. In this paper, we propose an extension of the encoder-decoder framework by adding a component called guiding network. The guiding network models the attribute properties of input images, and its output is leveraged to compose the input of the decoder at each time step. The guiding network can be plugged into the current encoder-decoder framework and trained in an end-to-end manner. Hence, the guiding vector can be adaptively learned according to the signal from the decoder, making itself to embed information from both image and language. Additionally, discriminative supervision can be employed to further improve the quality of guidance. The advantages of our proposed approach are verified by experiments carried out on the MS COCO dataset.

📄 PDF Abstract BibTeX arXiv:1804.00887

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeDecoderImage Captioning

Similar Papers 제목 키워드 기반

Fast Image Caption Generation with Position Alignment

2019-12-13 · Zheng-cong Fei

Recent neural network models for image captioning usually employ an encoder-decoder architecture, where the decoder adopts a recursive sequence decoding way. However, such autoregressive decoding may result in sequential…

Caption GenerationDecoderImage CaptioningPosition+1

AGIC: Attention-Guided Image Captioning to Improve Caption Relevance

2025-08-09 · L. D. M. S. Sai Teja, Ashok Urlana, Pruthwik Mishra arxiv

Despite significant progress in image captioning, generating accurate and descriptive captions remains a long-standing challenge. In this study, we propose Attention-Guided Image Captioning (AGIC), which amplifies salien…

Image Captioning

Guiding Image Captioning Models Toward More Specific Captions

2023-07-31 · ICCV 2023 1 · Simon Kornblith, Lala Li, ZiRui Wang, Thao Nguyen

Image captioning is conventionally formulated as the task of generating captions for images that match the distribution of reference image-caption pairs. However, reference captions in standard captioning datasets are sh…

Image CaptioningImage Retrieval

Transferable Decoding with Visual Entities for Zero-Shot Image Captioning

2023-07-31 · ICCV 2023 1 · Junjie Fei, Teng Wang, Jinrui Zhang, Zhenyu He 외

Image-to-text generation aims to describe images using natural language. Recently, zero-shot image captioning based on pre-trained vision-language models (VLMs) and large language models (LLMs) has made significant progr…

Caption GenerationHallucinationImage CaptioningImage to text+2

Brain Captioning: Decoding human brain activity into images and text

2023-05-19 · Matteo Ferrante, Furkan Ozcelik, Tommaso Boccato, Rufin VanRullen 외

Every day, the human brain processes an immense volume of visual information, relying on intricate neural mechanisms to perceive and interpret these stimuli. Recent breakthroughs in functional magnetic resonance imaging …

Brain DecodingDepth EstimationImage CaptioningImage Reconstruction+2