paper-with-me

Papers

Investigating Metric Diversity for Evaluating Long Document Summarisation

2022-10-01 · sdp (COLING) 2022 10 · Cai Yang, Stephen Wan

Long document summarisation, a challenging summarisation scenario, is the focus of the recently proposed LongSumm shared task. One of the limitations of this shared task has been its use of a single family of metrics for evaluation (the ROUGE metrics). In contrast, other fields, like text generation, employ multiple metrics. We replicated the LongSumm evaluation using multiple test set samples (vs. the single test set of the official shared task) and investigated how different metrics might complement each other in this evaluation framework. We show that under this more rigorous evaluation, (1) some of the key learnings from Longsumm 2020 and 2021 still hold, but the relative ranking of systems changes, and (2) the use of additional metrics reveals additional high-quality summaries missed by ROUGE, and (3) we show that SPICE is a candidate metric for summarisation evaluation for LongSumm.

📄 PDF Abstract BibTeX

Code (1)

caiyangcy/sdp-longsumm-metric-diversity 공식 구현 pytorch

Tasks

DiversityText Generation

Similar Papers 제목 키워드 기반

LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation

2023-09-21 · Jennifer A Bishop, Qianqian Xie, Sophia Ananiadou

Maintaining factual consistency is a critical issue in abstractive text summarisation, however, it cannot be assessed by traditional automatic metrics used for evaluating text summarisation, such as ROUGE scoring. Recent…

Improved Evidence Extraction and Metrics for Document Inconsistency Detection with LLMs

2026-01-06 · Nelvin Tan, Yaowen Zhang, James Asikin Cheung, Fusheng Liu 외 arxiv

Large language models (LLMs) are becoming useful in many domains due to their impressive abilities that arise from large training datasets and large model sizes. However, research on LLM-based approaches to document inco…

CoverageBench: Evaluating Information Coverage across Tasks and Domains

2026-03-20 · Saron Samuel, Andrew Yates, Dawn Lawrie, Ian Soboroff 외 arxiv

We wish to measure the information coverage of an ad hoc retrieval algorithm, that is, how much of the range of available relevant information is covered by the search results. Information coverage is a central aspect fo…

How Far are We from Robust Long Abstractive Summarization?

2022-10-30 · Huan Yee Koh, Jiaxin Ju, He Zhang, Ming Liu 외

Abstractive summarization has made tremendous progress in recent years. In this work, we perform fine-grained human annotations to evaluate long document abstractive summarization systems (i.e., models and metrics) with …

Abstractive Text Summarization

From Single to Multi: How LLMs Hallucinate in Multi-Document Summarization

2024-10-17 · Catarina G. Belem, Pouya Pezeskhpour, Hayate Iso, Seiji Maekawa 외

Although many studies have investigated and reduced hallucinations in large language models (LLMs) for single-document tasks, research on hallucination in multi-document summarization (MDS) tasks remains largely unexplor…

Document SummarizationHallucinationMulti-Document Summarization