paper-with-me

홈 › Papers

ContextRef: Evaluating Referenceless Metrics For Image Description Generation

2023-09-21 · Elisa Kreiss, Eric Zelikman, Christopher Potts, Nick Haber

Referenceless metrics (e.g., CLIPScore) use pretrained vision--language models to assess image descriptions directly without costly ground-truth reference texts. Such methods can facilitate rapid progress, but only if they truly align with human preference judgments. In this paper, we introduce ContextRef, a benchmark for assessing referenceless metrics for such alignment. ContextRef has two components: human ratings along a variety of established quality dimensions, and ten diverse robustness checks designed to uncover fundamental weaknesses. A crucial aspect of ContextRef is that images and descriptions are presented in context, reflecting prior work showing that context is important for description quality. Using ContextRef, we assess a variety of pretrained models, scoring functions, and techniques for incorporating context. None of the methods is successful with ContextRef, but we show that careful fine-tuning yields substantial improvements. ContextRef remains a challenging benchmark though, in large part due to the challenge of context dependence.

📄 PDF Abstract BibTeX arXiv:2309.11710

Code (1)

elisakreiss/contextref 공식 구현 pytorch

Tasks

Image Description

Methods 이 논문이 사용한 방법론

None 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Context Matters for Image Descriptions for Accessibility: Challenges for Referenceless Evaluation Metrics

2022-05-21 · Elisa Kreiss, Cynthia Bennett, Shayan Hooshmand, Eric Zelikman 외

Few images on the Web receive alt-text descriptions that would make them accessible to blind and low vision (BLV) users. Image-based NLG systems have progressed to the point where they can begin to address this persisten…

NoRefER: a Referenceless Quality Metric for Automatic Speech Recognition via Semi-Supervised Language Model Fine-Tuning with Contrastive Learning

2023-06-21 · Kamer Ali Yuksel, Thiago Ferreira, Golara Javadi, Mohamed El-Badrashiny 외

This paper introduces NoRefER, a novel referenceless quality metric for automatic speech recognition (ASR) systems. Traditional reference-based metrics for evaluating ASR systems require costly ground-truth transcripts. …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive LearningLanguage Modeling+4

Image quality assessment for determining efficacy and limitations of Super-Resolution Convolutional Neural Network (SRCNN)

2019-05-14 · Chris M. Ward, Josh Harguess, Brendan Crabb, Shibin Parameswaran

Traditional metrics for evaluating the efficacy of image processing techniques do not lend themselves to understanding the capabilities and limitations of modern image processing methods - particularly those enabled by d…

Image Quality AssessmentSSIMSuper-Resolution

ContextRefine-CLIP for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2025

2025-06-12 · Jing He, YiQing Wang, Lingling Li, Kexin Zhang 외

This report presents ContextRefine-CLIP (CR-CLIP), an efficient model for visual-textual multi-instance retrieval tasks. The approach is based on the dual-encoder AVION, on which we introduce a cross-modal attention flow…

Cross-Modal RetrievalEnsemble LearningMulti-Instance RetrievalRetrieval

Cross-validating Image Description Datasets and Evaluation Metrics

2016-05-01 · LREC 2016 5 · Josiah Wang, Robert Gaizauskas

The task of automatically generating sentential descriptions of image content has become increasingly popular in recent years, resulting in the development of large-scale image description datasets and the proposal of va…

Image DescriptionSentence