paper-with-me

홈 › Papers

JWSign: A Highly Multilingual Corpus of Bible Translations for more Diversity in Sign Language Processing

2023-11-16 · Shester Gueuwou, Sophie Siake, Colin Leong, Mathias Müller

Advancements in sign language processing have been hindered by a lack of sufficient data, impeding progress in recognition, translation, and production tasks. The absence of comprehensive sign language datasets across the world's sign languages has widened the gap in this field, resulting in a few sign languages being studied more than others, making this research area extremely skewed mostly towards sign languages from high-income countries. In this work we introduce a new large and highly multilingual dataset for sign language translation: JWSign. The dataset consists of 2,530 hours of Bible translations in 98 sign languages, featuring more than 1,500 individual signers. On this dataset, we report neural machine translation experiments. Apart from bilingual baseline systems, we also train multilingual systems, including some that take into account the typological relatedness of signed or spoken languages. Our experiments highlight that multilingual systems are superior to bilingual baselines, and that in higher-resource scenarios, clustering language pairs that are related improves translation quality.

📄 PDF Abstract BibTeX arXiv:2311.10174

Code (1)

shesterg/jwsign-machine-translation 공식 구현

Tasks

DiversityMachine TranslationSign Language TranslationTranslation

Similar Papers 제목 키워드 기반

The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration

2020-05-01 · LREC 2020 5 · Arya D. McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller 외

We present findings from the creation of a massively parallel corpus in over 1600 languages, the Johns Hopkins University Bible Corpus (JHUBC). The corpus consists of over 4000 unique translations of the Christian Bible …

Creating a massively parallel Bible corpus

2014-05-01 · LREC 2014 5 · Thomas Mayer, Michael Cysouw

We present our ongoing effort to create a massively parallel Bible corpus. While an ever-increasing number of Bible translations is available in electronic form on the internet, there is no large-scale parallel Bible cor…

Machine TranslationTranslation

Efficacy of ByT5 in Multilingual Translation of Biblical Texts for Underrepresented Languages

2024-05-22 · Corinne Aars, Lauren Adams, Xiaokan Tian, Zhaoyu Wang 외

This study presents the development and evaluation of a ByT5-based multilingual translation model tailored for translating the Bible into underrepresented languages. Utilizing the comprehensive Johns Hopkins University B…

The EDGeS Diachronic Bible Corpus

2020-05-01 · LREC 2020 5 · Gerlof Bouma, Evie Couss{\'e}, Trude Dijkstra, Nicoline van der Sijs

We present the EDGeS Diachronic Bible Corpus: a diachronically and synchronically parallel corpus of Bible translations in Dutch, English, German and Swedish, with texts from the 14th century until today. It is compiled …

The eBible Corpus: Data and Model Benchmarks for Bible Translation for Low-Resource Languages

2023-04-19 · Vesa Akerman, David Baines, Damien Daspit, Ulf Hermjakob 외

Efficiently and accurately translating a corpus into a low-resource language remains a challenge, regardless of the strategies employed, whether manual, automated, or a combination of the two. Many Christian organization…

BenchmarkingMachine TranslationNMTTranslation