paper-with-me

Papers

Introducing the Welsh Text Summarisation Dataset and Baseline Systems

2022-05-05 · LREC 2022 6 · Ignatius Ezeani, Mahmoud El-Haj, Jonathan Morris, Dawn Knight

Welsh is an official language in Wales and is spoken by an estimated 884,300 people (29.2% of the population of Wales). Despite this status and estimated increase in speaker numbers since the last (2011) census, Welsh remains a minority language undergoing revitalization and promotion by Welsh Government and relevant stakeholders. As part of the effort to increase the availability of Welsh digital technology, this paper introduces the first Welsh summarisation dataset, which we provide freely for research purposes to help advance the work on Welsh text summarization. The dataset was created by Welsh speakers by manually summarising Welsh Wikipedia articles. In addition, the paper discusses the implementation and evaluation of different summarisation systems for Welsh. The summarization systems and results will serve as benchmarks for the development of summarises in other minority language contexts.

📄 PDF Abstract BibTeX arXiv:2205.02545

Code (1)

ucrel/welsh-summarization-dataset 공식 구현

Tasks

ArticlesText Summarization

Similar Papers 제목 키워드 기반

Creation of an Evaluation Corpus and Baseline Evaluation Scores for Welsh Text Summarisation

2022-06-01 · CLTW (LREC) 2022 6 · Mahmoud El-Haj, Ignatius Ezeani, Jonathan Morris, Dawn Knight

As part of the effort to increase the availability of Welsh digital technology, this paper introduces the first human vs metrics Welsh summarisation evaluation results and dataset, which we provide freely for research pu…

BU-TTS: An Open-Source, Bilingual Welsh-English, Text-to-Speech Corpus

2022-06-01 · CLTW (LREC) 2022 6 · Stephen Russell, Dewi Jones, Delyth Prys

This paper presents the design, collection and verification of a bilingual text-to-speech synthesis corpus for Welsh and English. The ever expanding voice collection currently contains almost 10 hours of recordings from …

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

A Supervised Approach to Extractive Summarisation of Scientific Papers

2017-06-13 · CONLL 2017 8 · Ed Collins, Isabelle Augenstein, Sebastian Riedel

Automatic summarisation is a popular approach to reduce a document to its main arguments. Recent research in the area has focused on neural approaches to summarisation, which can be very data-hungry. However, few large d…

Sentence

DaNewsroom: A Large-scale Danish Summarisation Dataset

2020-05-01 · LREC 2020 5 · Daniel Varab, Natalie Schluter

Dataset development for automatic summarisation systems is notoriously English-oriented. In this paper we present the first large-scale non-English language dataset specifically curated for automatic summarisation. The d…

Abstractive Text SummarizationArticles

Terrain signatures in Welsh settlement names

2026-08-27 · Oktay Karakuş, Can Eyupoglu arxiv

Landscapes are named, but whether names retain measurable environmental information beyond broad geographic structure is rarely tested. We analysed 3,757 Welsh settlements using a frozen, source-audited 24-element lexica…