paper-with-me

홈 › Papers

Understanding the Extent to which Content Quality Metrics Measure the Information Quality of Summaries

2021-11-01 · CoNLL (EMNLP) 2021 11 · Daniel Deutsch, Dan Roth

Reference-based metrics such as ROUGE or BERTScore evaluate the content quality of a summary by comparing the summary to a reference. Ideally, this comparison should measure the summary’s information quality by calculating how much information the summaries have in common. In this work, we analyze the token alignments used by ROUGE and BERTScore to compare summaries and argue that their scores largely cannot be interpreted as measuring information overlap. Rather, they are better estimates of the extent to which the summaries discuss the same topics. Further, we provide evidence that this result holds true for many other summarization evaluation metrics. The consequence of this result is that the most frequently used summarization evaluation metrics do not align with the community’s research goal, to generate summaries with high-quality information. However, we conclude by demonstrating that a recently proposed metric, QAEval, which scores summaries using question-answering, appears to better capture information quality than current evaluations, highlighting a direction for future research.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Understanding the Extent to which Summarization Evaluation Metrics Measure the Information Quality of Summaries

2020-10-23 · Daniel Deutsch, Dan Roth

Reference-based metrics such as ROUGE or BERTScore evaluate the content quality of a summary by comparing the summary to a reference. Ideally, this comparison should measure the summary's information quality by calculati…

Can We Predict The Human Preference For Text-to-Image Content Prior To Generation And Is It Even Useful To Do So?

2026-06-03 · Joong Ho Kim, Keith G. Mills arxiv

Diffusion Models (DM) have revolutionized text-driven generation by enabling the synthesis of high-quality, photorealistic visual content from user prompts. Whereas prior advances in visual generation such as VAEs and GA…

Evaluating Scholarly Impact: Towards Content-Aware Bibliometrics

2021-11-01 · EMNLP 2021 11 · Saurav Manchanda, George Karypis

Quantitatively measuring the impact-related aspects of scientific, engineering, and technological (SET) innovations is a fundamental problem with broad applications. Traditional citation-based measures for assessing the …

ISO/IEC JTC1/SC29/WG11 MPEG2014/ m34661: Quality Assessment of High Dynamic Range (HDR) Video Content Using Existing Full-Reference Metrics

2018-03-13

The main focus of this document is to evaluate the performance of the existing LDR and HDR metrics on HDR video content which in turn will allow for a better understanding of how well each of these metrics work and if th…

Comparing Fair Ranking Metrics

2020-09-02 · Amifa Raj, Michael D. Ekstrand

Ranked lists are frequently used by information retrieval (IR) systems to present results believed to be relevant to the users information need. Fairness is a relatively new but important aspect of these rankings to meas…

FairnessInformation RetrievalRecommendation SystemsRetrieval