Generating captions without looking beyond objects
This paper explores new evaluation perspectives for image captioning and introduces a noun translation task that achieves comparative image caption generation performance by translating from a set of nouns to captions. This implies that in image captioning, all word categories other than nouns can be evoked by a powerful language model without sacrificing performance on n-gram precision. The paper also investigates lower and upper bounds of how much individual word categories in the captions contribute to the final BLEU score. A large possible improvement exists for nouns, verbs, and prepositions.
Code (0)
등록된 구현이 없습니다.
Tasks
Caption GenerationImage CaptioningLanguage ModelingLanguage ModellingTranslationSimilar Papers 제목 키워드 기반
Meaning guided video captioning
Current video captioning approaches often suffer from problems of missing objects in the video to be described, while generating captions semantically similar with ground truth sentences. In this paper, we propose a new …
Decoderobject-detectionObject DetectionVideo CaptioningWhat is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text Inputs
Given an input image, and nothing else, our method returns the bounding boxes of objects in the image and phrases that describe the objects. This is achieved within an open world paradigm, in which the objects in the inp…
BenchmarkingImage CaptioningImage to textPhrase Grounding+2Beyond Caption To Narrative: Video Captioning With Multiple Sentences
Recent advances in image captioning task have led to increasing interests in video captioning task. However, most works on video captioning are focused on generating single input of aggregated features, which hardly devi…
Action LocalizationImage CaptioningVideo CaptioningMOC-GAN: Mixing Objects and Captions to Generate Realistic Images
Generating images with conditional descriptions gains increasing interests in recent years. However, existing conditional inputs are suffering from either unstructured forms (captions) or limited information and expensiv…
Implicit RelationsRethinking the Reference-based Distinctive Image Captioning
Distinctive Image Captioning (DIC) -- generating distinctive captions that describe the unique details of a target image -- has received considerable attention over the last few years. A recent DIC work proposes to gener…
AttributeBenchmarkingImage Captioning