Geo-Aware Image Caption Generation
Standard image caption generation systems produce generic descriptions of images and do not utilize any contextual information or world knowledge. In particular, they are unable to generate captions that contain references to the geographic context of an image, for example, the location where a photograph is taken or relevant geographic objects around an image location. In this paper, we develop a geo-aware image caption generation system, which incorporates geographic contextual information into a standard image captioning pipeline. We propose a way to build an image-specific representation of the geographic context and adapt the caption generation network to produce appropriate geographic names in the image descriptions. We evaluate our system on a novel captioning dataset that contains contextualized captions and geographic metadata and achieve substantial improvements in BLEU, ROUGE, METEOR and CIDEr scores. We also introduce a new metric to assess generated geographic references directly and empirically demonstrate our system{'}s ability to produce captions with relevant and factually accurate geographic referencing.
Code (0)
등록된 구현이 없습니다.
Tasks
Caption GenerationImage CaptioningWorld KnowledgeSimilar Papers 제목 키워드 기반
Transform, Contrast and Tell: Coherent Entity-Aware Multi-Image Captioning
Coherent entity-aware multi-image captioning aims to generate coherent captions for neighboring images in a news document. There are coherence relationships among neighboring images because they often describe same entit…
Caption GenerationCoherence EvaluationContrastive LearningImage CaptioningJournalistic Guidelines Aware News Image Captioning
The task of news article image captioning aims to generate descriptive and informative captions for news article images. Unlike conventional image captions that simply describe the content of the image in general terms, …
Caption GenerationDescriptiveImage CaptioningCOSMic: A Coherence-Aware Generation Metric for Image Descriptions
Developers of text generation models rely on automated evaluation metrics as a stand-in for slow and expensive manual evaluations. However, image captioning metrics have struggled to give accurate learned estimates of th…
Caption GenerationImage CaptioningText GenerationTemporal Knowledge-Aware Image Captioning
Contextualized image captioning is a task that extends beyond generating a purely visual description of the image content and aims to produce a caption that is influenced by the context and informed by the real world kno…
Caption GenerationImage CaptioningWorld KnowledgeCIC: A Framework for Culturally-Aware Image Captioning
Image Captioning generates descriptive sentences from images using Vision-Language Pre-trained models (VLPs) such as BLIP, which has improved greatly. However, current methods lack the generation of detailed descriptive …
DescriptiveImage CaptioningQuestion AnsweringVisual Question Answering+1