End-to-end Image Captioning Exploits Multimodal Distributional Similarity
We hypothesize that end-to-end neural image captioning systems work seemingly
well because they exploit and learn distributional similarity' in a multimodal
feature space by mapping a test image to similar training images in this space
and generating a caption from the same space. To validate our hypothesis, we
focus on the image' side of image captioning, and vary the input image
representation but keep the RNN text generation component of a CNN-RNN model
constant. Our analysis indicates that image captioning models (i) are capable
of separating structure from noisy input representations; (ii) suffer virtually
no significant performance loss when a high dimensional representation is
compressed to a lower dimensional space; (iii) cluster images with similar
visual and linguistic information together. Our findings indicate that our
distributional similarity hypothesis holds. We conclude that regardless of the
image representation used image captioning systems seem to match images and
generate captions in a learned joint image-text semantic subspace.
Code (0)
등록된 구현이 없습니다.
Tasks
Image CaptioningText GenerationSimilar Papers 제목 키워드 기반
End-to-end Image Captioning Exploits Distributional Similarity in Multimodal Space
We hypothesize that end-to-end neural image captioning systems work seemingly well because they exploit and learn {`}distributional similarity{'} in a multimodal feature space, by mapping a test image to similar training…
Image CaptioningText GenerationWhat is image captioning made of?
We hypothesize that end-to-end neural image captioning systems work seemingly well because they exploit and learn ‘distributional similarity’ in a multimodal feature space, by mapping a test image to similar training ima…
Image CaptioningText GenerationCaptioning Images with Diverse Objects
Recent captioning models are limited in their ability to scale and describe concepts unseen in paired image-text corpora. We propose the Novel Object Captioner (NOC), a deep visual semantic captioning model that can desc…
ObjectObject RecognitioniPIC-XAI: Improving PIC-XAI for Enhanced Image Captioning Explanation
Image captioning task with its complexity has taken advantage of the recent developments in Deep learning (DL). However, DL-models are fundamentally abstruse and explaining their behaviour is a challenge. In this paper w…
Image CaptioningTAGGoing Beneath the Surface: Evaluating Image Captioning for Grammaticality, Truthfulness and Diversity
Image captioning as a multimodal task has drawn much interest in recent years. However, evaluation for this task remains a challenging problem. Existing evaluation metrics focus on surface similarity between a candidate …
DiagnosticDiversityImage Captioning