paper-with-me

홈 › Papers

Evaluation Under Imperfect Benchmarks and Ratings: A Case Study in Text Simplification

2025-04-13 · Joseph Liu, Yoonsoo Nam, Xinyue Cui, Swabha Swayamdipta

Despite the successes of language models, their evaluation remains a daunting challenge for new and existing tasks. We consider the task of text simplification, commonly used to improve information accessibility, where evaluation faces two major challenges. First, the data in existing benchmarks might not reflect the capabilities of current language models on the task, often containing disfluent, incoherent, or simplistic examples. Second, existing human ratings associated with the benchmarks often contain a high degree of disagreement, resulting in inconsistent ratings; nevertheless, existing metrics still have to show higher correlations with these imperfect ratings. As a result, evaluation for the task is not reliable and does not reflect expected trends (e.g., more powerful models being assigned higher scores). We address these challenges for the task of text simplification through three contributions. First, we introduce SynthSimpliEval, a synthetic benchmark for text simplification featuring simplified sentences generated by models of varying sizes. Through a pilot study, we show that human ratings on our benchmark exhibit high inter-annotator agreement and reflect the expected trend: larger models produce higher-quality simplifications. Second, we show that auto-evaluation with a panel of LLM judges (LLMs-as-a-jury) often suffices to obtain consistent ratings for the evaluation of text simplification. Third, we demonstrate that existing learnable metrics for text simplification benefit from training on our LLMs-as-a-jury-rated synthetic data, closing the gap with pure LLMs-as-a-jury for evaluation. Overall, through our case study on text simplification, we show that a reliable evaluation requires higher quality test data, which could be obtained through synthetic data and LLMs-as-a-jury ratings.

📄 PDF Abstract BibTeX arXiv:2504.09394

Code (0)

등록된 구현이 없습니다.

Tasks

Text Simplification

Similar Papers 제목 키워드 기반

Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas

2025-09-26 · Luke Guerdan, Justin Whitehouse, Kimberly Truong, Kenneth Holstein 외 arxiv

As Generative AI (GenAI) systems see growing adoption, a key concern involves the external validity of evaluations, or the extent to which they generalize from lab-based to real-world deployment conditions. Threats to th…

Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling

2025-04-18 · Shaomu Tan, Christof Monz

A key challenge in MT evaluation is the inherent noise and inconsistency of human ratings. Regression-based neural metrics struggle with this noise, while prompting LLMs shows promise at system-level evaluation but perfo…

Machine TranslationTranslation

Spontaneous Reward Hacking in Iterative Self-Refinement

2024-07-05 · Jane Pan, He He, Samuel R. Bowman, Shi Feng

Language models are capable of iteratively improving their outputs based on natural language feedback, thus enabling in-context optimization of user preference. In place of human users, a second language model can be use…

Language ModelingLanguage Modelling

Compare without Despair: Reliable Preference Evaluation with Generation Separability

2024-07-02 · Sayan Ghosh, Tejas Srinivasan, Swabha Swayamdipta

Human evaluation of generated language through pairwise preference judgments is pervasive. However, under common scenarios, such as when generations from a model pair are very similar, or when stochastic decoding results…

AMR4NLI: Interpretable and robust NLI measures from semantic graphs

2023-06-01 · Juri Opitz, Shira Wein, Julius Steen, Anette Frank 외

The task of natural language inference (NLI) asks whether a given premise (expressed in NL) entails a given NL hypothesis. NLI benchmarks contain human ratings of entailment, but the meaning relationships driving these r…

Natural Language InferenceSentence