Studying Summarization Evaluation Metrics in the Appropriate Scoring Range
In summarization, automatic evaluation metrics are usually compared based on their ability to correlate with human judgments. Unfortunately, the few existing human judgment datasets have been created as by-products of the manual evaluations performed during the DUC/TAC shared tasks. However, modern systems are typically better than the best systems submitted at the time of these shared tasks. We show that, surprisingly, evaluation metrics which behave similarly on these datasets (average-scoring range) strongly disagree in the higher-scoring range in which current systems now operate. It is problematic because metrics disagree yet we can{'}t decide which one to trust. This is a call for collecting human judgments for high-scoring summaries as this would resolve the debate over which metrics to trust. This would also be greatly beneficial to further improve summarization systems and metrics alike.
Code (1)
Similar Papers 제목 키워드 기반
Metrics also Disagree in the Low Scoring Range: Revisiting Summarization Evaluation Metrics
In text summarization, evaluating the efficacy of automatic metrics without human judgments has become recently popular. One exemplar work concludes that automatic metrics strongly disagree when ranking high-scoring summ…
Text SummarizationRe-evaluating Evaluation in Text Summarization
Automated evaluation metrics as a stand-in for manual evaluation are an essential part of the development of text-generation tasks such as text summarization. However, while the field has progressed, our standard metrics…
Text GenerationText SummarizationLearning to Score System Summaries for Better Content Selection Evaluation.
The evaluation of summaries is a challenging but crucial task of the summarization field. In this work, we propose to learn an automatic scoring metric based on the human judgements available as part of classical summari…
Document SummarizationMulti-Document SummarizationSemantic Textual SimilarityHuman-like Summarization Evaluation with ChatGPT
Evaluating text summarization is a challenging problem, and existing evaluation metrics are far from satisfactory. In this study, we explored ChatGPT's ability to perform human-like summarization evaluation using four hu…
Text SummarizationCAWESumm: A Contextual and Anonymous Walk Embedding Based Extractive Summarization of Legal Bills
Extractive summarization of lengthy legal documents requires an appropriate sentence scoring mechanism. This mechanism should capture both the local semantics of a sentence as well as the global document-level context of…
Document EmbeddingExtractive SummarizationSentenceSentence Embedding+1