paper-with-me

Papers

HOLMS: Alternative Summary Evaluation with Large Language Models

2020-12-01 · COLING 2020 8 · Yassine Mrabet, Dina Demner-Fushman

Efficient document summarization requires evaluation measures that can not only rank a set of systems based on an average score, but also highlight which individual summary is better than another. However, despite the very active research on summarization approaches, few works have proposed new evaluation measures in the recent years. The standard measures relied upon for the development of summarization systems are most often ROUGE and BLEU which, despite being efficient in overall system ranking, remain lexical in nature and have a limited potential when it comes to training neural networks. In this paper, we present a new hybrid evaluation measure for summarization, called HOLMS, that combines both language models pre-trained on large corpora and lexical similarity measures. Through several experiments, we show that HOLMS outperforms ROUGE and BLEU substantially in its correlation with human judgments on several extractive summarization datasets for both linguistic quality and pyramid scores.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Document SummarizationExtractive Summarization

Similar Papers 제목 키워드 기반

Bornholmsk Natural Language Processing: Resources and Tools

2019-09-01 · WS (NoDaLiDa) 2019 9 · Leon Derczynski, Alex Speed Kjeldsen

This paper introduces language processing resources and tools for Bornholmsk, a language spoken on the island of Bornholm, with roots in Danish and closely related to Scanian. This presents an overview of the language an…

One Prompt To Rule Them All: LLMs for Opinion Summary Evaluation

2024-02-18 · Tejpalsingh Siledar, Swaroop Nath, Sankara Sri Raghava Ravindra Muddu, Rupasai Rangaraju 외

Evaluation of opinion summaries using conventional reference-based metrics rarely provides a holistic evaluation and has been shown to have a relatively low correlation with human judgments. Recent studies suggest using …

Allnlg evaluationOpinion SummarizationSpecificity

Data Augmentation for Low-Resource Dialogue Summarization

2022-01-16 · ACL ARR January 2022 1 · Anonymous

We present DADS, a novel Data Augmentation technique for low-resource Dialogue Summarization. Our method generates synthetic examples by replacing sections of text from both the input dialogue and summary while preservin…

Data AugmentationMeeting Summarization

Data Augmentation for Low-Resource Dialogue Summarization

2022-07-01 · Findings (NAACL) 2022 7 · Yongtai Liu, Joshua Maynez, Gonçalo Simões, Shashi Narayan

We present DADS, a novel Data Augmentation technique for low-resource Dialogue Summarization. Our method generates synthetic examples by replacing sections of text from both the input dialogue and summary while preservin…

Data AugmentationMeeting Summarization

Moral Hazard in Multi-Agent Language Models

2026-07-27 · Dane Malenfant arxiv

Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmström's model of moral hazard in teams, we introduce the Dialogue Moral Hazard Game, a theo…