Quantifying the amount of visual information used by neural caption generators
This paper addresses the sensitivity of neural image caption generators to their visual input. A sensitivity analysis and omission analysis based on image foils is reported, showing that the extent to which image captioning architectures retain and are sensitive to visual information varies depending on the type of word being generated and the position in the caption as a whole. We motivate this work in the context of broader goals in the field to achieve more explainability in AI.
Code (1)
Tasks
Image CaptioningPositionSensitivitySimilar Papers 제목 키워드 기반
Visual News: Benchmark and Challenges in News Image Captioning
We propose Visual News Captioner, an entity-aware model for the task of news image captioning. We also introduce Visual News, a large-scale benchmark consisting of more than one million news images along with associated …
ArticlesImage CaptioningVideo Captioning: a comparative review of where we are and which could be the route
Video captioning is the process of describing the content of a sequence of images capturing its semantic relationships and meanings. Dealing with this task with a single image is arduous, not to mention how difficult it …
Video CaptioningOracle performance for visual captioning
The task of associating images and videos with a natural language description has attracted a great amount of attention recently. Rapid progress has been made in terms of both developing novel algorithms and releasing ne…
Image CaptioningLanguage ModelingLanguage ModellingVideo CaptioningVisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning
The ability to quickly learn from a small quantity oftraining data widens the range of machine learning applications. In this paper, we propose a data-efficient image captioning model, VisualGPT, which leverages the ling…
DecoderImage CaptioningLanguage ModellingMedical Report GenerationWatch, Listen and Tell: Multi-modal Weakly Supervised Dense Event Captioning
Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning…
Sound Source Localization