paper-with-me

홈 › Papers

Quantifying the amount of visual information used by neural caption generators

2018-10-12 · Marc Tanti, Albert Gatt, Kenneth P. Camilleri

This paper addresses the sensitivity of neural image caption generators to their visual input. A sensitivity analysis and omission analysis based on image foils is reported, showing that the extent to which image captioning architectures retain and are sensitive to visual information varies depending on the type of word being generated and the position in the caption as a whole. We motivate this work in the context of broader goals in the field to achieve more explainability in AI.

📄 PDF Abstract BibTeX arXiv:1810.05475

Code (1)

mtanti/quantifing-visual-information 공식 구현 tf

Tasks

Image CaptioningPositionSensitivity

Similar Papers 제목 키워드 기반

Visual News: Benchmark and Challenges in News Image Captioning

2020-10-08 · EMNLP 2021 11 · Fuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente Ordonez

We propose Visual News Captioner, an entity-aware model for the task of news image captioning. We also introduce Visual News, a large-scale benchmark consisting of more than one million news images along with associated …

ArticlesImage Captioning

Video Captioning: a comparative review of where we are and which could be the route

2022-04-12 · Daniela Moctezuma, Tania Ramírez-delReal, Guillermo Ruiz, Othón González-Chávez

Video captioning is the process of describing the content of a sequence of images capturing its semantic relationships and meanings. Dealing with this task with a single image is arduous, not to mention how difficult it …

Video Captioning

Oracle performance for visual captioning

2015-11-14 · Li Yao, Nicolas Ballas, Kyunghyun Cho, John R. Smith 외

The task of associating images and videos with a natural language description has attracted a great amount of attention recently. Rapid progress has been made in terms of both developing novel algorithms and releasing ne…

Image CaptioningLanguage ModelingLanguage ModellingVideo Captioning

VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning

2021-02-20 · CVPR 2022 1 · Jun Chen, Han Guo, Kai Yi, Boyang Li 외

The ability to quickly learn from a small quantity oftraining data widens the range of machine learning applications. In this paper, we propose a data-efficient image captioning model, VisualGPT, which leverages the ling…

DecoderImage CaptioningLanguage ModellingMedical Report Generation

Watch, Listen and Tell: Multi-modal Weakly Supervised Dense Event Captioning

2019-09-22 · ICCV 2019 10 · Tanzila Rahman, Bicheng Xu, Leonid Sigal

Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning…

Sound Source Localization