paper-with-me

홈 › Papers

On the use of human reference data for evaluating automatic image descriptions

2020-06-15 · Emiel van Miltenburg

Automatic image description systems are commonly trained and evaluated using crowdsourced, human-generated image descriptions. The best-performing system is then determined using some measure of similarity to the reference data (BLEU, Meteor, CIDER, etc). Thus, both the quality of the systems as well as the quality of the evaluation depends on the quality of the descriptions. As Section 2 will show, the quality of current image description datasets is insufficient. I argue that there is a need for more detailed guidelines that take into account the needs of visually impaired users, but also the feasibility of generating suitable descriptions. With high-quality data, evaluation of image description systems could use reference descriptions, but we should also look for alternatives.

📄 PDF Abstract BibTeX arXiv:2006.08792

Code (0)

등록된 구현이 없습니다.

Tasks

Image Description

Similar Papers 제목 키워드 기반

ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation

2023-04-12 · NeurIPS 2023 11 · Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong 외

We present a comprehensive solution to learn and improve text-to-image models from human preference feedback. To begin with, we build ImageReward -- the first general-purpose text-to-image human preference reward model -…

Image GenerationPreference MappingText to Image GenerationText-to-Image Generation

VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions

2025-09-30 · Kazuki Matsuda, Yuiga Wada, Shinnosuke Hirano, Seitaro Otsuki 외 arxiv

In this study, we focus on the automatic evaluation of long and detailed image captions generated by multimodal Large Language Models (MLLMs). Most existing automatic evaluation metrics for image captioning are primarily…

Image Captioning

Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation

2023-05-02 · NeurIPS 2023 11 · Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana 외

The ability to collect a large dataset of human preferences from text-to-image users is usually limited to companies, making such datasets inaccessible to the public. To address this issue, we create a web app that enabl…

Image GenerationPreference MappingText to Image GenerationText-to-Image Generation

Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference

2024-12-31 · Mingqi Gao, Yixin Liu, Xinyu Hu, Xiaojun Wan 외

Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences. Due to the high cost and time-consuming nature of human evaluations, an autom…

HP-Edit: A Human-Preference Post-Training Framework for Image Editing

2026-04-21 · Fan Li, Chonghuinan Wang, Lina Lei, Yuping Qiu 외 arxiv

Common image editing tasks typically adopt powerful generative diffusion models as the leading paradigm for real-world content editing. Meanwhile, although reinforcement learning (RL) methods such as Diffusion-DPO and Fl…

Reinforcement LearningImage Editing