ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus
With more than 7000 languages worldwide, multilingual natural language processing (NLP) is essential both from an academic and commercial perspective. Researching typological properties of languages is fundamental for progress in multilingual NLP. Examples include assessing language similarity for effective transfer learning, injecting inductive biases into machine learning models or creating resources such as dictionaries and inflection tables. We provide ParCourE, an online tool that allows to browse a word-aligned parallel corpus, covering 1334 languages. We give evidence that this is useful for typological research. ParCourE can be set up for any parallel corpus and can thus be used for typological research on other corpora as well as for exploring their quality and properties.
Code (0)
등록된 구현이 없습니다.
Tasks
Multilingual NLPTransfer LearningSimilar Papers 제목 키워드 기반
Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data
This paper investigates a critical design decision in the practice of massively multilingual continual pre-training -- the inclusion of parallel data. Specifically, we study the impact of bilingual translation data for m…
TranslationCVSS Corpus and Massively Multilingual Speech-to-Speech Translation
We introduce CVSS, a massively multilingual-to-English speech-to-speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English. CVSS is derived from the Common Voice speech …
SentenceSpeech-to-Speech TranslationSpeech-to-TextSpeech-to-Text Translation+1PMIndiaSum: Multilingual and Cross-lingual Headline Summarization for Languages in India
This paper introduces PMIndiaSum, a multilingual and massively parallel summarization corpus focused on languages in India. Our corpus provides a training and testing ground for four language families, 14 languages, and …
Cross-Lingual Abstractive SummarizationMultilingual NLPText SummarizationThe Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration
We present findings from the creation of a massively parallel corpus in over 1600 languages, the Johns Hopkins University Bible Corpus (JHUBC). The corpus consists of over 4000 unique translations of the Christian Bible …
nmT5 - Is parallel data still relevant for pre-training massively multilingual language models?
Recently, mT5 - a massively multilingual version of T5 - leveraged a unified text-to-text format to attain state-of-the-art results on a wide variety of multilingual NLP tasks. In this paper, we investigate the impact of…
Language ModelingLanguage ModellingMachine TranslationMultilingual NLP+1