TLAXCALA: a multilingual corpus of independent news
We acquire corpora from the domain of independent news from the Tlaxcala website. We build monolingual corpora for 15 languages and parallel corpora for all the combinations of those 15 languages. These corpora include languages for which only very limited such resources exist (e.g. Tamazight). We present the acquisition process in detail and we also present detailed statistics of the produced corpora, concerning mainly quantitative dimensions such as the size of the corpora per language (for the monolingual corpora) and per language pair (for the parallel corpora). To the best of our knowledge, these are the first publicly available parallel and monolingual corpora for the domain of independent news. We also create models for unsupervised sentence splitting for all the languages of the study.
Code (0)
등록된 구현이 없습니다.
Tasks
Language IdentificationMachine TranslationSentenceSimilar Papers 제목 키워드 기반
MultiNews: A Web collection of an Aligned Multimodal and Multilingual Corpus
Integrating Natural Language Processing (NLP) and computer vision is a promising effort. However, the applicability of these methods directly depends on the availability of a specific multimodal data that includes images…
ArticlesContent-Based Image RetrievalImage RetrievalMachine Translation+1Multilingual Open Text Release 1: Public Domain News in 44 Languages
We present Multilingual Open Text (MOT), a new multilingual corpus containing text in 44 languages, many of which have limited existing text resources for natural language processing. The first release of the corpus cont…
ArticlesA Multilingual Simplified Language News Corpus
Simplified language news articles are being offered by specialized web portals in several countries. The thousands of articles that have been published over the years are a valuable resource for natural language processi…
ArticlesText SimplificationBianet: A Parallel News Corpus in Turkish, Kurdish and English
We present a new open-source parallel corpus consisting of news articles collected from the Bianet magazine, an online newspaper that publishes Turkish news, often along with their translations in English and Kurdish. In…
ArticlesMachine TranslationTranslationEuronews: a multilingual speech corpus for ASR
In this paper we present a multilingual speech corpus, designed for Automatic Speech Recognition (ASR) purposes. Data come from the portal Euronews and were acquired both from the Web and from TV. The corpus includes dat…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+1