paper-with-me

Papers

Monolingual and Parallel Corpora for Kangri Low Resource Language

2021-03-22 · Shweta Chauhan, Shefali Saxena, Philemon Daniel

In this paper we present the dataset of Himachali low resource endangered language, Kangri (ISO 639-3xnr) listed in the United Nations Educational, Scientific and Cultural Organization (UNESCO). The compilation of kangri corpus has been a challenging task due to the non-availability of the digitalized resources. The corpus contains 1,81,552 Monolingual and 27,362 Hindi-Kangri Parallel corpora. We shared pre-trained kangri word embeddings. We also reported the Bilingual Evaluation Understudy (BLEU) score and Metric for Evaluation of Translation with Explicit ORdering (METEOR) score of Statistical Machine Translation (SMT) and Neural Machine Translation (NMT) results for the corpus. The corpus is freely available for non-commercial usages and research. To the best of our knowledge, this is the first Himachali low resource endangered language corpus. The resources are available at (https://github.com/chauhanshweta/Kangri_corpus)

📄 PDF Abstract BibTeX arXiv:2103.11596

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNMTTranslationWord Embeddings

Similar Papers 제목 키워드 기반

Zero-Resource Neural Machine Translation with Monolingual Pivot Data

2019-11-01 · WS 2019 11 · Anna Currey, Kenneth Heafield

Zero-shot neural machine translation (NMT) is a framework that uses source-pivot and target-pivot parallel data to train a source-target NMT system. An extension to zero-shot NMT is zero-resource NMT, which generates pse…

Machine TranslationNMTTranslation

TLAXCALA: a multilingual corpus of independent news

2014-05-01 · LREC 2014 5 · Antonio Toral

We acquire corpora from the domain of independent news from the Tlaxcala website. We build monolingual corpora for 15 languages and parallel corpora for all the combinations of those 15 languages. These corpora include l…

Language IdentificationMachine TranslationSentence

Semi-Supervised Learning for Neural Machine Translation

2016-06-15 · ACL 2016 8 · Yong Cheng, Wei Xu, Zhongjun He, wei he 외

While end-to-end neural machine translation (NMT) has made remarkable progress recently, NMT systems only rely on parallel corpora for parameter estimation. Since parallel corpora are usually limited in quantity, quality…

DecoderMachine TranslationNMTparameter estimation+1

ERNIE-M: Enhanced Multilingual Representation by Aligning Cross-lingual Semantics with Monolingual Corpora

2020-12-31 · EMNLP 2021 11 · Xuan Ouyang, Shuohuan Wang, Chao Pang, Yu Sun 외

Recent studies have demonstrated that pre-trained cross-lingual models achieve impressive performance in downstream cross-lingual tasks. This improvement benefits from learning a large amount of monolingual and parallel …

SentenceTranslation

MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages

2022-06-01 · EAMT 2022 6 · Marta Bañón, Miquel Esplà-Gomis, Mikel L. Forcada, Cristian García-Romero 외

We introduce the project “MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages”, funded by the Connecting Europe Facility, which is aimed at building monolingual a…