paper-with-me

홈 › Papers

LCEval: Learned Composite Metric for Caption Evaluation

2020-12-24 · Naeha Sharif, Lyndon White, Mohammed Bennamoun, Wei Liu, Syed Afaq Ali Shah

Automatic evaluation metrics hold a fundamental importance in the development and fine-grained analysis of captioning systems. While current evaluation metrics tend to achieve an acceptable correlation with human judgements at the system level, they fail to do so at the caption level. In this work, we propose a neural network-based learned metric to improve the caption-level caption evaluation. To get a deeper insight into the parameters which impact a learned metrics performance, this paper investigates the relationship between different linguistic features and the caption-level correlation of the learned metrics. We also compare metrics trained with different training examples to measure the variations in their evaluation. Moreover, we perform a robustness analysis, which highlights the sensitivity of learned and handcrafted metrics to various sentence perturbations. Our empirical analysis shows that our proposed metric not only outperforms the existing metrics in terms of caption-level correlation but it also shows a strong system-level correlation against human assessments.

📄 PDF Abstract BibTeX arXiv:2012.13136

Code (1)

NaehaSharif/LCEVal 공식 구현 tf

Tasks

Sentence

Similar Papers 제목 키워드 기반

Learning-based Composite Metrics for Improved Caption Evaluation

2018-07-01 · ACL 2018 7 · Naeha Sharif, Lyndon White, Mohammed Bennamoun, Syed Afaq Ali Shah

The evaluation of image caption quality is a challenging task, which requires the assessment of two main aspects in a caption: adequacy and fluency. These quality aspects can be judged using a combination of several ling…

Image CaptioningLanguage ModelingLanguage ModellingSemantic Textual Similarity+1

Polos: Multimodal Metric Learning from Human Feedback for Image Captioning

2024-02-28 · CVPR 2024 1 · Yuiga Wada, Kanta Kaneda, Daichi Saito, Komei Sugiura

Establishing an automatic evaluation metric that closely aligns with human judgments is essential for effectively developing image captioning models. Recent data-driven metrics have demonstrated a stronger correlation wi…

Contrastive LearningImage CaptioningMetric Learning

FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model

2024-06-10 · Yebin Lee, Imseong Park, Myungjoo Kang

Most existing image captioning evaluation metrics focus on assigning a single numerical score to a caption by comparing it with reference captions. However, these methods do not provide an explanation for the assigned sc…

Image Captioning

LLM-Free Image Captioning Evaluation in Reference-Flexible Settings

2025-12-25 · Shinnosuke Hirano, Yuiga Wada, Kazuki Matsuda, Seitaro Otsuki 외 arxiv

We focus on the automatic evaluation of image captions in both reference-based and reference-free settings. Existing metrics based on large language models (LLMs) favor their own generations; therefore, the neutrality is…

Image Captioning

DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning

2024-09-28 · Kazuki Matsuda, Yuiga Wada, Komei Sugiura

In this work, we address the challenge of developing automatic evaluation metrics for image captioning, with a particular focus on robustness against hallucinations. Existing metrics are often inadequate for handling hal…

HallucinationImage Captioning