paper-with-me

Papers

ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus

2021-07-14 · ACL 2021 5 · Ayyoob Imani, Masoud Jalili Sabet, Philipp Dufter, Michael Cysouw, Hinrich Schütze

With more than 7000 languages worldwide, multilingual natural language processing (NLP) is essential both from an academic and commercial perspective. Researching typological properties of languages is fundamental for progress in multilingual NLP. Examples include assessing language similarity for effective transfer learning, injecting inductive biases into machine learning models or creating resources such as dictionaries and inflection tables. We provide ParCourE, an online tool that allows to browse a word-aligned parallel corpus, covering 1334 languages. We give evidence that this is useful for typological research. ParCourE can be set up for any parallel corpus and can thus be used for typological research on other corpora as well as for exploring their quality and properties.

📄 PDF Abstract BibTeX arXiv:2107.06632

Code (0)

등록된 구현이 없습니다.

Tasks

Multilingual NLPTransfer Learning

Similar Papers 제목 키워드 기반

Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data

2025-05-31 · Shaoxiong Ji, Zihao Li, Jaakko Paavola, Indraneil Paul 외

This paper investigates a critical design decision in the practice of massively multilingual continual pre-training -- the inclusion of parallel data. Specifically, we study the impact of bilingual translation data for m…

Translation

CVSS Corpus and Massively Multilingual Speech-to-Speech Translation

2022-01-11 · LREC 2022 6 · Ye Jia, Michelle Tadmor Ramanovich, Quan Wang, Heiga Zen

We introduce CVSS, a massively multilingual-to-English speech-to-speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English. CVSS is derived from the Common Voice speech …

SentenceSpeech-to-Speech TranslationSpeech-to-TextSpeech-to-Text Translation+1

PMIndiaSum: Multilingual and Cross-lingual Headline Summarization for Languages in India

2023-05-15 · Ashok Urlana, Pinzhen Chen, Zheng Zhao, Shay B. Cohen 외

This paper introduces PMIndiaSum, a multilingual and massively parallel summarization corpus focused on languages in India. Our corpus provides a training and testing ground for four language families, 14 languages, and …

Cross-Lingual Abstractive SummarizationMultilingual NLPText Summarization

The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration

2020-05-01 · LREC 2020 5 · Arya D. McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller 외

We present findings from the creation of a massively parallel corpus in over 1600 languages, the Johns Hopkins University Bible Corpus (JHUBC). The corpus consists of over 4000 unique translations of the Christian Bible …

nmT5 - Is parallel data still relevant for pre-training massively multilingual language models?

2021-08-01 · ACL 2021 5 · Mihir Kale, Aditya Siddhant, Rami Al-Rfou, Linting Xue 외

Recently, mT5 - a massively multilingual version of T5 - leveraged a unified text-to-text format to attain state-of-the-art results on a wide variety of multilingual NLP tasks. In this paper, we investigate the impact of…

Language ModelingLanguage ModellingMachine TranslationMultilingual NLP+1