Going Beneath the Surface: Evaluating Image Captioning for Grammaticality, Truthfulness and Diversity
Image captioning as a multimodal task has drawn much interest in recent years. However, evaluation for this task remains a challenging problem. Existing evaluation metrics focus on surface similarity between a candidate caption and a set of reference captions, and do not check the actual relation between a caption and the underlying visual content. We introduce a new diagnostic evaluation framework for the task of image captioning, with the goal of directly assessing models for grammaticality, truthfulness and diversity (GTD) of generated captions. We demonstrate the potential of our evaluation framework by evaluating existing image captioning models on a wide ranging set of synthetic datasets that we construct for diagnostic evaluation. We empirically show how the GTD evaluation framework, in combination with diagnostic datasets, can provide insights into model capabilities and limitations to supplement standard evaluations.
Code (0)
등록된 구현이 없습니다.
Tasks
DiagnosticDiversityImage CaptioningSimilar Papers 제목 키워드 기반
Capacitive Sensor Based 2D Subsurface Imaging Technology for Non Destructive Evaluation of Building Surfaces
Understanding the underlying structure of building surfaces like walls and floors is essential when carrying out building maintenance and modification work. To facilitate such work, this paper introduces a capacitive sen…
Connotation Lexicon: A Dash of Sentiment Beneath the Surface Meaning
AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages
Multimodal AI research has overwhelmingly focused on high-resource languages, hindering the democratization of advancements in the field. To address this, we present AfriCaption, a comprehensive framework for multilingua…
Image CaptioningImaging dynamics beneath turbid media via parallelized single-photon detection
Noninvasive optical imaging through dynamic scattering media has numerous important biomedical applications but still remains a challenging task. While standard diffuse imaging methods measure optical absorption or fluor…
Video ReconstructionExperimenting with Self-Supervision using Rotation Prediction for Image Captioning
Image captioning is a task in the field of Artificial Intelligence that merges between computer vision and natural language processing. It is responsible for generating legends that describe images, and has various appli…
DecoderImage CaptioningSelf-Supervised Learning