paper-with-me

Papers

LR-Sum: Summarization for Less-Resourced Languages

2022-12-19 · Chester Palen-Michel, Constantine Lignos

This preprint describes work in progress on LR-Sum, a new permissively-licensed dataset created with the goal of enabling further research in automatic summarization for less-resourced languages. LR-Sum contains human-written summaries for 40 languages, many of which are less-resourced. We describe our process for extracting and filtering the dataset from the Multilingual Open Text corpus (Palen-Michel et al., 2022). The source data is public domain newswire collected from from Voice of America websites, and LR-Sum is released under a Creative Commons license (CC BY 4.0), making it one of the most openly-licensed multilingual summarization datasets. We describe how we plan to use the data for modeling experiments and discuss limitations of the dataset.

📄 PDF Abstract BibTeX arXiv:2212.09674

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Comparing Approaches to Automatic Summarization in Less-Resourced Languages

2025-12-30 · Chester Palen-Michel, Constantine Lignos arxiv

Automatic text summarization has achieved high performance in high-resourced languages like English, but comparatively less attention has been given to summarization in less-resourced languages. This work compares a vari…

Text SummarizationData Augmentation

QFS-Composer: Query-focused summarization pipeline for less resourced languages

2026-04-12 · Vuk Đuranović, Marko Robnik Šikonja arxiv

Large language models (LLMs) demonstrate strong performance in text summarization, yet their effectiveness drops significantly across languages with restricted training resources. This work addresses the challenge of que…

Question GenerationText SummarizationQuestion Answering

Unsupervised Approach to Multilingual User Comments Summarization

2021-04-01 · EACL (Hackashop) 2021 4 · Aleš Žagar, Marko Robnik-Šikonja

User commenting is a valuable feature of many news outlets, enabling them a contact with readers and enabling readers to express their opinion, provide different viewpoints, and even complementary information. Yet, large…

Extractive SummarizationSentence

Sequence to sequence pretraining for a less-resourced Slovenian language

2022-07-28 · Matej Ulčar, Marko Robnik-Šikonja

Large pretrained language models have recently conquered the area of natural language processing. As an alternative to predominant masked language modelling introduced in BERT, the T5 model has introduced a more general …

Language ModelingLanguage ModellingMachine TranslationOpen-Domain Question Answering+4

Open or Closed LLM for Lesser-Resourced Languages? Lessons from Greek

2025-01-22 · John Pavlopoulos, Juli Bakagianni, Kanella Pouli, Maria Gavriilidou

Natural Language Processing (NLP) for lesser-resourced languages faces persistent challenges, including limited datasets, inherited biases from high-resource languages, and the need for domain-specific solutions. This st…

Authorship Attribution