paper-with-me

홈 › Papers

Gradations of Error Severity in Automatic Image Descriptions

2020-12-01 · INLG (ACL) 2020 12 · Emiel van Miltenburg, Wei-Ting Lu, Emiel Krahmer, Albert Gatt, Guanyi Chen, Lin Li, Kees Van Deemter

Earlier research has shown that evaluation metrics based on textual similarity (e.g., BLEU, CIDEr, Meteor) do not correlate well with human evaluation scores for automatically generated text. We carried out an experiment with Chinese speakers, where we systematically manipulated image descriptions to contain different kinds of errors. Because our manipulated descriptions form minimal pairs with the reference descriptions, we are able to assess the impact of different kinds of errors on the perceived quality of the descriptions. Our results show that different kinds of errors elicit significantly different evaluation scores, even though all erroneous descriptions differ in only one character from the reference descriptions. Evaluation metrics based solely on textual similarity are unable to capture these differences, which (at least partially) explains their poor correlation with human judgments. Our work provides the foundations for future work, where we aim to understand why different errors are seen as more or less severe.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Causally Grounded Taxonomy for Image Degradation Robustness Evaluation

2026-05-15 · Stefan Becker, Simon Weiss, Wolfgang Hübner, Michael Arens arxiv

Image degradations can occur during acquisition, processing, and transmission, altering visual appearance and affecting downstream vision tasks. They are studied in several communities, including synthetic corruption ben…

Image Quality Assessment

Do Large Language Models Judge Error Severity Like Humans?

2025-06-05 · Diege Sun, Guanyi Chen, Zhao Fan, Xiaorong Cheng 외

Large Language Models (LLMs) are increasingly used as automated evaluators in natural language generation, yet it remains unclear whether they can accurately replicate human judgments of error severity. In this study, we…

Text Generation

Always Clear Days: Degradation Type and Severity Aware All-In-One Adverse Weather Removal

2023-10-27 · Yu-Wei Chen, Soo-Chang Pei

All-in-one adverse weather removal is an emerging topic on image restoration, which aims to restore multiple weather degradations in an unified model, and the challenge are twofold. First, discover and handle the propert…

AllDomain AdaptationImage Restoration

ControlFusion: A Controllable Image Fusion Framework with Language-Vision Degradation Prompts

2025-03-30 · Linfeng Tang, Yeda Wang, Zhanchuan Cai, Junjun Jiang 외

Current image fusion methods struggle to address the composite degradations encountered in real-world imaging scenarios and lack the flexibility to accommodate user-specific requirements. In response to these challenges,…

MedQ-Deg: A Multidimensional Benchmark for Evaluating MLLMs Across Medical Image Quality Degradations

2026-03-08 · Jiyao Liu, Junzhi Ning, Chenglong Ma, Wanying Qu 외 arxiv

Despite impressive performance on standard benchmarks, multimodal large language models (MLLMs) face critical challenges in real-world clinical environments where medical images inevitably suffer various quality degradat…