paper-with-me

홈 › Papers

Metrics also Disagree in the Low Scoring Range: Revisiting Summarization Evaluation Metrics

2020-11-08 · COLING 2020 8 · Manik Bhandari, Pranav Gour, Atabak Ashfaq, PengFei Liu

In text summarization, evaluating the efficacy of automatic metrics without human judgments has become recently popular. One exemplar work concludes that automatic metrics strongly disagree when ranking high-scoring summaries. In this paper, we revisit their experiments and find that their observations stem from the fact that metrics disagree in ranking summaries from any narrow scoring range. We hypothesize that this may be because summaries are similar to each other in a narrow scoring range and are thus, difficult to rank. Apart from the width of the scoring range of summaries, we analyze three other properties that impact inter-metric agreement - Ease of Summarization, Abstractiveness, and Coverage. To encourage reproducible research, we make all our analysis code and data publicly available.

📄 PDF Abstract BibTeX arXiv:2011.04096

Code (0)

등록된 구현이 없습니다.

Tasks

Text Summarization

Similar Papers 제목 키워드 기반

Studying Summarization Evaluation Metrics in the Appropriate Scoring Range

2019-07-01 · ACL 2019 7 · Maxime Peyrard

In summarization, automatic evaluation metrics are usually compared based on their ability to correlate with human judgments. Unfortunately, the few existing human judgment datasets have been created as by-products of th…

Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering

2025-11-10 · Sai Shridhar Balamurali, Lu Cheng arxiv

Evaluating answers from state-of-the-art large language models (LLMs) is challenging: lexical metrics miss semantic nuances, whereas "LLM-as-Judge" scoring is computationally expensive. We re-evaluate a lightweight alter…

Natural Language InferenceQuestion Answering

A Kernel Score Perspective on Forecast Disagreement and the Linear Pool

2024-12-12 · Fabian Krüger

The variance of a linearly combined forecast distribution (or linear pool) consists of two components: The average variance of the component distributions (`average uncertainty'), and the average squared difference betwe…

Revisiting the Task of Scoring Open IE Relations

2018-05-01 · LREC 2018 5 · William L{\'e}chelle, Philippe Langlais
Knowledge Base CompletionLanguage ModelingLanguage ModellingOpen Information Extraction

Revisiting the Role of Similarity and Dissimilarity in Best Counter Argument Retrieval

2023-04-18 · Hongguang Shi, Shuirong Cao, Cam-Tu Nguyen

This paper studies the task of best counter-argument retrieval given an input argument. Following the definition that the best counter-argument addresses the same aspects as the input argument while having the opposite s…

Argument RetrievalLearning-To-RankRetrieval