paper-with-me

홈 › Papers

Evaluating Multiple System Summary Lengths: A Case Study

2018-10-01 · EMNLP 2018 10 · Ori Shapira, David Gabay, Hadar Ronen, Judit Bar-Ilan, Yael Amsterdamer, Ani Nenkova, Ido Dagan

Practical summarization systems are expected to produce summaries of varying lengths, per user needs. While a couple of early summarization benchmarks tested systems across multiple summary lengths, this practice was mostly abandoned due to the assumed cost of producing reference summaries of multiple lengths. In this paper, we raise the research question of whether reference summaries of a single length can be used to reliably evaluate system summaries of multiple lengths. For that, we have analyzed a couple of datasets as a case study, using several variants of the ROUGE metric that are standard in summarization evaluation. Our findings indicate that the evaluation protocol in question is indeed competitive. This result paves the way to practically evaluating varying-length summaries with simple, possibly existing, summarization benchmarks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LCFO: Long Context and Long Form Output Dataset and Benchmarking

2024-12-11 · Marta R. Costa-jussà, Pierre Andrews, Mariano Coria Meglioli, Joy Chen 외

This paper presents the Long Context and Form Output (LCFO) benchmark, a novel evaluation framework for assessing gradual summarization and summary expansion capabilities across diverse domains. LCFO consists of long inp…

BenchmarkingForm

How to Find Strong Summary Coherence Measures? A Toolbox and a Comparative Study for Summary Coherence Measure Evaluation

2022-09-14 · COLING 2022 10 · Julius Steen, Katja Markert

Automatically evaluating the coherence of summaries is of great significance both to enable cost-efficient summarizer evaluation and as a tool for improving coherence by selecting high-scoring candidate summaries. While …

What comes next? Extractive summarization by next-sentence prediction

2019-01-12 · Jingyun Liu, Jackie C. K. Cheung, Annie Louis

Existing approaches to automatic summarization assume that a length limit for the summary is given, and view content selection as an optimization problem to maximize informativeness and minimize redundancy within this bu…

Extractive SummarizationInformativenessPredictionSentence

A Large-Scale Multi-Length Headline Corpus for Analyzing Length-Constrained Headline Generation Model Evaluation

2019-03-28 · WS 2019 10 · Yuta Hitomi, Yuya Taguchi, Hideaki Tamori, Ko Kikuta 외

Browsing news articles on multiple devices is now possible. The lengths of news article headlines have precise upper bounds, dictated by the size of the display of the relevant device or interface. Therefore, controlling…

ArticlesHeadline Generation

Abstractiveness Metrics for Evaluating Text Summarization: A Refined Formulation with Empirical Validation

2026-07-12 · Praveenkumar Katwe, Rakesh Chandra Balabantaray, Kali Prasad Vittala arxiv

Quantifying abstractiveness in generated summaries is essential for evaluating summarization models beyond surface-level metrics like ROUGE. We introduce Reference Abstraction (RA), Summary Abstraction (SA), and Abstract…

Text Summarization