paper-with-me

홈 › Papers

VIFIDEL: Evaluating the Visual Fidelity of Image Descriptions

2019-07-22 · ACL 2019 7 · Pranava Madhyastha, Josiah Wang, Lucia Specia

We address the task of evaluating image description generation systems. We propose a novel image-aware metric for this task: VIFIDEL. It estimates the faithfulness of a generated caption with respect to the content of the actual image, based on the semantic similarity between labels of objects depicted in images and words in the description. The metric is also able to take into account the relative importance of objects mentioned in human reference descriptions during evaluation. Even if these human reference descriptions are not available, VIFIDEL can still reliably evaluate system descriptions. The metric achieves high correlation with human judgments on two well-known datasets and is competitive with metrics that depend on human references

📄 PDF Abstract BibTeX arXiv:1907.09340

Code (0)

등록된 구현이 없습니다.

Tasks

Image DescriptionSemantic SimilaritySemantic Textual Similarity

Similar Papers 제목 키워드 기반

On the use of human reference data for evaluating automatic image descriptions

2020-06-15 · Emiel van Miltenburg

Automatic image description systems are commonly trained and evaluated using crowdsourced, human-generated image descriptions. The best-performing system is then determined using some measure of similarity to the referen…

Image Description

Can Unified Generation and Understanding Models Maintain Semantic Equivalence Across Different Output Modalities?

2026-02-27 · Hongbo Jiang, Jie Li, Yunhang Shen, Pingyang Dai 외 arxiv

Unified Multimodal Large Language Models (U-MLLMs) integrate understanding and generation within a single architecture. However, existing evaluations typically assess these capabilities separately, overlooking semantic e…

ImageInWords: Unlocking Hyper-Detailed Image Descriptions

2024-05-05 · Roopal Garg, Andrea Burns, Burcu Karagol Ayan, Yonatan Bitton 외

Despite the longstanding adage "an image is worth a thousand words," generating accurate hyper-detailed image descriptions remains unsolved. Trained on short web-scraped image text, vision-language models often generate …

Image GenerationSpecificityText to Image GenerationText-to-Image Generation

CAP: Evaluation of Persuasive and Creative Image Generation

2024-12-10 · Aysan Aghazadeh, Adriana Kovashka

We address the task of advertisement image generation and introduce three evaluation metrics to assess Creativity, prompt Alignment, and Persuasiveness (CAP) in generated advertisement images. Despite recent advancements…

Image GenerationPersuasiveness

UFC-BERT: Unifying Multi-Modal Controls for Conditional Image Synthesis

2021-05-21 · NeurIPS 2021 12 · Zhu Zhang, Jianxin Ma, Chang Zhou, Rui Men 외

Conditional image synthesis aims to create an image according to some multi-modal guidance in the forms of textual descriptions, reference images, and image blocks to preserve, as well as their combinations. In this pape…

Image Generation