paper-with-me

Papers

Creating a massively parallel Bible corpus

2014-05-01 · LREC 2014 5 · Thomas Mayer, Michael Cysouw

We present our ongoing effort to create a massively parallel Bible corpus. While an ever-increasing number of Bible translations is available in electronic form on the internet, there is no large-scale parallel Bible corpus that allows language researchers to easily get access to the texts and their parallel structure for a large variety of different languages. We report on the current status of the corpus, with over 900 translations in more than 830 language varieties. All translations are tokenized (e.g., separating punctuation marks) and Unicode normalized. Mainly due to copyright restrictions only portions of the texts are made publicly available. However, we provide co-occurrence information for each translation in a (sparse) matrix format. All word forms in the translation are given together with their frequency and the verses in which they occur.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration

2020-05-01 · LREC 2020 5 · Arya D. McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller 외

We present findings from the creation of a massively parallel corpus in over 1600 languages, the Johns Hopkins University Bible Corpus (JHUBC). The corpus consists of over 4000 unique translations of the Christian Bible …

The EDGeS Diachronic Bible Corpus

2020-05-01 · LREC 2020 5 · Gerlof Bouma, Evie Couss{\'e}, Trude Dijkstra, Nicoline van der Sijs

We present the EDGeS Diachronic Bible Corpus: a diachronically and synchronically parallel corpus of Bible translations in Dutch, English, German and Swedish, with texts from the 14th century until today. It is compiled …

Towards a Broad Coverage Named Entity Resource: A Data-Efficient Approach for Many Diverse Languages

2022-01-28 · LREC 2022 6 · Silvia Severini, Ayyoob Imani, Philipp Dufter, Hinrich Schütze

Parallel corpora are ideal for extracting a multilingual named entity (MNE) resource, i.e., a dataset of names translated into multiple languages. Prior work on extracting MNE datasets from parallel corpora required reso…

Bilingual Lexicon InductionTransliteration

Deriving Consensus for Multi-Parallel Corpora: an English Bible Study

2017-11-01 · IJCNLP 2017 11 · Patrick Xia, David Yarowsky

What can you do with multiple noisy versions of the same text? We present a method which generates a single consensus between multi-parallel corpora. By maximizing a function of linguistic features between word pairs, we…

Machine Translation

Evaluating prose style transfer with the Bible

2017-11-13 · Keith Carlson, Allen Riddell, Daniel Rockmore

In the prose style transfer task a system, provided with text input and a target prose style, produces output which preserves the meaning of the input text but alters the style. These systems require parallel data for ev…

Style Transfer