paper-with-me

홈 › Papers

Do Large Language Models Judge Error Severity Like Humans?

2025-06-05 · Diege Sun, Guanyi Chen, Zhao Fan, Xiaorong Cheng, Tingting He

Large Language Models (LLMs) are increasingly used as automated evaluators in natural language generation, yet it remains unclear whether they can accurately replicate human judgments of error severity. In this study, we systematically compare human and LLM assessments of image descriptions containing controlled semantic errors. We extend the experimental framework of van Miltenburg et al. (2020) to both unimodal (text-only) and multimodal (text + image) settings, evaluating four error types: age, gender, clothing type, and clothing colour. Our findings reveal that humans assign varying levels of severity to different error types, with visual context significantly amplifying perceived severity for colour and type errors. Notably, most LLMs assign low scores to gender errors but disproportionately high scores to colour errors, unlike humans, who judge both as highly severe but for different reasons. This suggests that these models may have internalised social norms influencing gender judgments but lack the perceptual grounding to emulate human sensitivity to colour, which is shaped by distinct neural mechanisms. Only one of the evaluated LLMs, Doubao, replicates the human-like ranking of error severity, but it fails to distinguish between error types as clearly as humans. Surprisingly, DeepSeek-V3, a unimodal LLM, achieves the highest alignment with human judgments across both unimodal and multimodal conditions, outperforming even state-of-the-art multimodal models.

📄 PDF Abstract BibTeX arXiv:2506.05142

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Not All Errors are Equal: Learning Text Generation Metrics using Stratified Error Synthesis

2022-10-10 · Wenda Xu, YiLin Tuan, Yujie Lu, Michael Saxon 외

Is it possible to build a general and automatic natural language generation (NLG) evaluation metric? Existing learned metrics either perform unsatisfactorily or are restricted to tasks where large human rating data is al…

AllImage CaptioningMachine Translationnlg evaluation+1

ERRORQUAKE: Heavy-Tailed Error Severity Distributions in Open-Weight Large Language Models

2026-04-15 · Jason Z Wang arxiv

At matched accuracy, open-weight LLMs differ substantially in the shape of their error severity distribution -- a difference invisible to the scalar error rate. Hallucination benchmarks report a single error count and tr…

FineRadScore: A Radiology Report Line-by-Line Evaluation Technique Generating Corrections with Severity Scores

2024-05-31 · Alyssa Huang, Oishi Banerjee, Kay Wu, Eduardo Pontes Reis 외

The current gold standard for evaluating generated chest x-ray (CXR) reports is through radiologist annotations. However, this process can be extremely time-consuming and costly, especially when evaluating large numbers …

Language ModelingLanguage ModellingLarge Language Model

CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation

2026-03-06 · Mohammed Baharoon, Thibault Heintz, Siavash Raissi, Mahmoud Alabbad 외 arxiv

We introduce CRIMSON, a clinically grounded evaluation framework for chest X-ray report generation that assesses reports based on diagnostic correctness, contextual relevance, and patient safety. Unlike prior metrics, CR…

From Outliers to Errors: Auditing Pali-to-English LLM Translations with Multi-Reference Adjudication

2026-05-31 · Máté Metzger, Nadnapang Phophichit, Hansa Dhammahaso arxiv

Single-score translation metrics can conflate legitimate variation with error, a problem especially acute for classical languages where multiple defensible English renderings of the same passage coexist. We audit Pali-to…