End-to-end Image Captioning Exploits Distributional Similarity in Multimodal Space
We hypothesize that end-to-end neural image captioning systems work seemingly well because they exploit and learn {}distributional similarity{'} in a multimodal feature space, by mapping a test image to similar training images in this space and generating a caption from the same space. To validate our hypothesis, we focus on the {}image{'} side of image captioning, and vary the input image representation but keep the RNN text generation model of a CNN-RNN constant. Our analysis indicates that image captioning models (i) are capable of separating structure from noisy input representations; (ii) experience virtually no significant performance loss when a high dimensional representation is compressed to a lower dimensional space; (iii) cluster images with similar visual and linguistic information together. Our experiments all point to one fact: that our distributional similarity hypothesis holds. We conclude that, regardless of the image representation, image captioning systems seem to match images and generate captions in a learned joint image-text semantic subspace.
Code (1)
Tasks
Image CaptioningText GenerationSimilar Papers 제목 키워드 기반
End-to-end Image Captioning Exploits Multimodal Distributional Similarity
We hypothesize that end-to-end neural image captioning systems work seemingly well because they exploit and learn `distributional similarity' in a multimodal feature space by mapping a test image to similar training imag…
Image CaptioningText GenerationWhat is image captioning made of?
We hypothesize that end-to-end neural image captioning systems work seemingly well because they exploit and learn ‘distributional similarity’ in a multimodal feature space, by mapping a test image to similar training ima…
Image CaptioningText GenerationCaptioning Images with Diverse Objects
Recent captioning models are limited in their ability to scale and describe concepts unseen in paired image-text corpora. We propose the Novel Object Captioner (NOC), a deep visual semantic captioning model that can desc…
ObjectObject RecognitioniPIC-XAI: Improving PIC-XAI for Enhanced Image Captioning Explanation
Image captioning task with its complexity has taken advantage of the recent developments in Deep learning (DL). However, DL-models are fundamentally abstruse and explaining their behaviour is a challenge. In this paper w…
Image CaptioningTAGGoing Beneath the Surface: Evaluating Image Captioning for Grammaticality, Truthfulness and Diversity
Image captioning as a multimodal task has drawn much interest in recent years. However, evaluation for this task remains a challenging problem. Existing evaluation metrics focus on surface similarity between a candidate …
DiagnosticDiversityImage Captioning