paper-with-me

Papers

JW300: A Wide-Coverage Parallel Corpus for Low-Resource Languages

2019-07-01 · ACL 2019 7 · {\v{Z}}eljko Agi{\'c}, Ivan Vuli{\'c}

Viable cross-lingual transfer critically depends on the availability of parallel texts. Shortage of such resources imposes a development and evaluation bottleneck in multilingual processing. We introduce JW300, a parallel corpus of over 300 languages with around 100 thousand parallel sentences per language pair on average. In this paper, we present the resource and showcase its utility in experiments with cross-lingual word embedding induction and multi-source part-of-speech projection.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual Transfer

Similar Papers 제목 키워드 기반

Towards a Broad Coverage Named Entity Resource: A Data-Efficient Approach for Many Diverse Languages

2022-01-28 · LREC 2022 6 · Silvia Severini, Ayyoob Imani, Philipp Dufter, Hinrich Schütze

Parallel corpora are ideal for extracting a multilingual named entity (MNE) resource, i.e., a dataset of names translated into multiple languages. Prior work on extracting MNE datasets from parallel corpora required reso…

Bilingual Lexicon InductionTransliteration

From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora

2025-05-20 · Yingli Shen, Wen Lai, Shuo Wang, Kangyang Luo 외

Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However, the unaligned nature of such data limi…

GlotLID: Language Identification for Low-Resource Languages

2023-10-24 · Amir Hossein Kargaran, Ayyoob Imani, François Yvon, Hinrich Schütze

Several recent papers have published good solutions for language identification (LID) for about 300 high-resource and medium-resource languages. However, there is no LID available that (i) covers a wide range of low-reso…

Dialect IdentificationLanguage Identification

ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus

2021-07-14 · ACL 2021 5 · Ayyoob Imani, Masoud Jalili Sabet, Philipp Dufter, Michael Cysouw 외

With more than 7000 languages worldwide, multilingual natural language processing (NLP) is essential both from an academic and commercial perspective. Researching typological properties of languages is fundamental for pr…

Multilingual NLPTransfer Learning

Bilingual Words and Phrase Mappings for Marathi and Hindi SMT

2017-10-05 · Sreelekha. S, Pushpak Bhattacharyya

Lack of proper linguistic resources is the major challenges faced by the Machine Translation system developments when dealing with the resource poor languages. In this paper, we describe effective ways to utilize the lex…

Machine TranslationTranslation