paper-with-me

홈 › Papers

A Multilingual Simplified Language News Corpus

2022-06-01 · READI (LREC) 2022 6 · Renate Hauser, Jannis Vamvas, Sarah Ebling, Martin Volk

Simplified language news articles are being offered by specialized web portals in several countries. The thousands of articles that have been published over the years are a valuable resource for natural language processing, especially for efforts towards automatic text simplification. In this paper, we present SNIML, a large multilingual corpus of news in simplified language. The corpus contains 13k simplified news articles written in one of six languages: Finnish, French, Italian, Swedish, English, and German. All articles are shared under open licenses that permit academic use. The level of text simplification varies depending on the news portal. We believe that even though SNIML is not a parallel corpus, it can be useful as a complement to the more homogeneous but often smaller corpora of news in the simplified variety of one language that are currently in use.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesText Simplification

Similar Papers 제목 키워드 기반

MultiNews: A Web collection of an Aligned Multimodal and Multilingual Corpus

2017-11-01 · WS 2017 11 · Haithem Afli, Pintu Lohar, Andy Way

Integrating Natural Language Processing (NLP) and computer vision is a promising effort. However, the applicability of these methods directly depends on the availability of a specific multimodal data that includes images…

ArticlesContent-Based Image RetrievalImage RetrievalMachine Translation+1

Multilingual Open Text Release 1: Public Domain News in 44 Languages

2022-01-14 · LREC 2022 6 · Chester Palen-Michel, June Kim, Constantine Lignos

We present Multilingual Open Text (MOT), a new multilingual corpus containing text in 44 languages, many of which have limited existing text resources for natural language processing. The first release of the corpus cont…

Articles

A Corpus for Sentence-level Subjectivity Detection on English News Articles

2023-05-29 · Francesco Antici, Andrea Galassi, Federico Ruggeri, Katerina Korre 외

We develop novel annotation guidelines for sentence-level subjectivity detection, which are not limited to language-specific cues. We use our guidelines to collect NewsSD-ENG, a corpus of 638 objective and 411 subjective…

ArticlesMachine TranslationSentence

Euronews: a multilingual speech corpus for ASR

2014-05-01 · LREC 2014 5 · Roberto Gretter

In this paper we present a multilingual speech corpus, designed for Automatic Speech Recognition (ASR) purposes. Data come from the portal Euronews and were acquired both from the Web and from TV. The corpus includes dat…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+1

Sina at FigNews 2024: Multilingual Datasets Annotated with Bias and Propaganda

2024-07-12 · Lina Duaibes, Areej Jaber, Mustafa Jarrar, Ahmad Qadi 외

The proliferation of bias and propaganda on social media is an increasingly significant concern, leading to the development of techniques for automatic detection. This article presents a multilingual corpus of 12, 000 Fa…