paper-with-me

Papers

DACSA: A large-scale Dataset for Automatic summarization of Catalan and Spanish newspaper Articles

2022-07-01 · NAACL 2022 7 · Encarnación Segarra Soriano, Vicent Ahuir, Lluís-F. Hurtado, José González

The application of supervised methods to automatic summarization requires the availability of adequate corpora consisting of a set of document-summary pairs. As in most Natural Language Processing tasks, the great majority of available datasets for summarization are in English, making it difficult to develop automatic summarization models for other languages. Although Spanish is gradually forming part of some recent summarization corpora, it is not the same for minority languages such as Catalan.In this work, we describe the construction of a corpus of Catalan and Spanish newspapers, the Dataset for Automatic summarization of Catalan and Spanish newspaper Articles (DACSA) corpus. It is a high-quality large-scale corpus that can be used to train summarization models for Catalan and Spanish.We have carried out an analysis of the corpus, both in terms of the style of the summaries and the difficulty of the summarization task. In particular, we have used a set of well-known metrics in the summarization field in order to characterize the corpus. Additionally, for benchmarking purposes, we have evaluated the performances of some extractive and abstractive summarization systems on the DACSA corpus.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Abstractive Text SummarizationArticlesBenchmarking

Similar Papers 제목 키워드 기반

CNewSum: A Large-scale Chinese News Summarization Dataset with Human-annotated Adequacy and Deducibility Level

2021-10-21 · Danqing Wang, Jiaze Chen, Xianze Wu, Hao Zhou 외

Automatic text summarization aims to produce a brief but crucial summary for the input documents. Both extractive and abstractive methods have witnessed great success in English datasets in recent years. However, there h…

News SummarizationText Summarization

LANS: Large-scale Arabic News Summarization Corpus

2022-10-24 · Abdulaziz Alhamadani, Xuchao Zhang, Jianfeng He, Chang-Tien Lu

Text summarization has been intensively studied in many languages, and some languages have reached advanced stages. Yet, Arabic Text Summarization (ATS) is still in its developing stages. Existing ATS datasets are either…

ArticlesDiversityNews SummarizationText Summarization

Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?

2025-03-21 · Jeremy Barnes, Naiara Perez, Alba Bonet-Jover, Begoña Altuna

Studies on evaluation metrics and LLM-as-a-Judge models for automatic text summarization have largely been focused on English, limiting our understanding of their effectiveness in other languages. Through our new dataset…

ArticlesText Summarization

Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model

2019-06-04 · ACL 2019 7 · Alexander R. Fabbri, Irene Li, Tianwei She, Suyi Li 외

Automatic generation of summaries from multiple news articles is a valuable tool as the number of online publications grows rapidly. Single document summarization (SDS) systems have benefited from advances in neural enco…

ArticlesDecoderDocument SummarizationExtractive Summarization+1

HowSumm: A Multi-Document Summarization Dataset Derived from WikiHow Articles

2021-10-07 · Odellia Boni, Guy Feigenblat, Guy Lev, Michal Shmueli-Scheuer 외

We present HowSumm, a novel large-scale dataset for the task of query-focused multi-document summarization (qMDS), which targets the use-case of generating actionable instructions from a set of sources. This use-case is …

Abstractive Text SummarizationArticlesDocument SummarizationMulti-Document Summarization