paper-with-me

홈 › Papers

IceSum: An Icelandic Text Summarization Corpus

2021-06-01 · NAACL 2021 4 · J{\'o}n Da{\dh}ason, Hrafn Loftsson, Salome Sigur{\dh}ard{\'o}ttir, {\TH}orsteinn Bj{\"o}rnsson

Automatic Text Summarization (ATS) is the task of generating concise and fluent summaries from one or more documents. In this paper, we present IceSum, the first Icelandic corpus annotated with human-generated summaries. IceSum consists of 1,000 online news articles and their extractive summaries. We train and evaluate several neural network-based models on this dataset, comparing them against a selection of baseline methods. We find that an encoder-decoder model with a sequence-to-sequence based extractor obtains the best results, outperforming all baseline methods. Furthermore, we evaluate how the size of the training corpus affects the quality of the generated summaries. We release the corpus and the models with an open license.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesDecoderText Summarization

Similar Papers 제목 키워드 기반

Compiling and Filtering ParIce: An English-Icelandic Parallel Corpus

2019-09-01 · WS (NoDaLiDa) 2019 9 · Starkaður Barkarson, Steinþór Steingrímsson

We present ParIce, a new English-Icelandic parallel corpus. This is the first parallel corpus built for the purposes of language technology development and research for Icelandic, although some Icelandic texts can be fou…

Icelandic Parallel Abstracts Corpus

2021-08-11 · Haukur Barri Símonarson, Vésteinn Snæbjarnarson

We present a new Icelandic-English parallel corpus, the Icelandic Parallel Abstracts Corpus (IPAC), composed of abstracts from student theses and dissertations. The texts were collected from the Skemman repository which …

NMTSentenceTranslation

A Warm Start and a Clean Crawled Corpus -- A Recipe for Good Language Models

2022-01-14 · Vésteinn Snæbjarnarson, Haukur Barri Símonarson, Pétur Orri Ragnarsson, Svanhvít Lilja Ingólfsdóttir 외

We train several language models for Icelandic, including IceBERT, that achieve state-of-the-art performance in a variety of downstream tasks, including part-of-speech tagging, named entity recognition, grammatical error…

Constituency ParsingGrammatical Error Detectionnamed-entity-recognitionNamed Entity Recognition+3

A Warm Start and a Clean Crawled Corpus - A Recipe for Good Language Models

2022-06-01 · LREC 2022 6 · Vésteinn Snæbjarnarson, Haukur Barri Símonarson, Pétur Orri Ragnarsson, Svanhvít Lilja Ingólfsdóttir 외

We train several language models for Icelandic, including IceBERT, that achieve state-of-the-art performance in a variety of downstream tasks, including part-of-speech tagging, named entity recognition, grammatical error…

Constituency ParsingGrammatical Error Detectionnamed-entity-recognitionNamed Entity Recognition+3

Creating Data in Icelandic for Text Normalization

2021-05-01 · NoDaLiDa 2021 5 · Helga Svala Sigurðardóttir, Anna Björk Nikulásdóttir, Jón Guðnason

There is no natural way to acquire normalized data so we try to create good enough data to attempt more advanced methods for text normalization. We manually annotated the first normalized corpus in Icelandic, 40,000 sent…

Text Normalization