VIFIDEL: Evaluating the Visual Fidelity of Image Descriptions
We address the task of evaluating image description generation systems. We propose a novel image-aware metric for this task: VIFIDEL. It estimates the faithfulness of a generated caption with respect to the content of the actual image, based on the semantic similarity between labels of objects depicted in images and words in the description. The metric is also able to take into account the relative importance of objects mentioned in human reference descriptions during evaluation. Even if these human reference descriptions are not available, VIFIDEL can still reliably evaluate system descriptions. The metric achieves high correlation with human judgments on two well-known datasets and is competitive with metrics that depend on human references
Code (0)
등록된 구현이 없습니다.
Tasks
Image DescriptionSemantic SimilaritySemantic Textual SimilaritySimilar Papers 제목 키워드 기반
On the use of human reference data for evaluating automatic image descriptions
Automatic image description systems are commonly trained and evaluated using crowdsourced, human-generated image descriptions. The best-performing system is then determined using some measure of similarity to the referen…
Image DescriptionCan Unified Generation and Understanding Models Maintain Semantic Equivalence Across Different Output Modalities?
Unified Multimodal Large Language Models (U-MLLMs) integrate understanding and generation within a single architecture. However, existing evaluations typically assess these capabilities separately, overlooking semantic e…
ImageInWords: Unlocking Hyper-Detailed Image Descriptions
Despite the longstanding adage "an image is worth a thousand words," generating accurate hyper-detailed image descriptions remains unsolved. Trained on short web-scraped image text, vision-language models often generate …
Image GenerationSpecificityText to Image GenerationText-to-Image GenerationCAP: Evaluation of Persuasive and Creative Image Generation
We address the task of advertisement image generation and introduce three evaluation metrics to assess Creativity, prompt Alignment, and Persuasiveness (CAP) in generated advertisement images. Despite recent advancements…
Image GenerationPersuasivenessUFC-BERT: Unifying Multi-Modal Controls for Conditional Image Synthesis
Conditional image synthesis aims to create an image according to some multi-modal guidance in the forms of textual descriptions, reference images, and image blocks to preserve, as well as their combinations. In this pape…
Image Generation