paper-with-me

홈 › Papers

EM Corpus: a comparable corpus for a less-resourced language pair Manipuri-English

2021-09-01 · RANLP (BUCC) 2021 9 · Rudali Huidrom, Yves Lepage, Khogendra Khomdram

In this paper, we introduce a sentence-level comparable text corpus crawled and created for the less-resourced language pair, Manipuri(mni) and English (eng). Our monolingual corpora comprise 1.88 million Manipuri sentences and 1.45 million English sentences, and our parallel corpus comprises 124,975 Manipuri-English sentence pairs. These data were crawled and collected over a year from August 2020 to March 2021 from a local newspaper website called ‘The Sangai Express.’ The resources reported in this paper are made available to help the low-resourced languages community for MT/NLP tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Projection of Argumentative Corpora from Source to Target Languages

2017-09-01 · WS 2017 9 · Ahmet Aker, Huangpan Zhang

Argumentative corpora are costly to create and are available in only few languages with English dominating the area. In this paper we release the first publicly available Mandarin argumentative corpus. The corpus is crea…

Argument MiningMachine TranslationTranslation

Developing a Fine-Grained Corpus for a Less-resourced Language: the case of Kurdish

2019-09-25 · WS 2019 8 · Roshna Omer Abdulrahman, Hossein Hassani, Sina Ahmadi

Kurdish is a less-resourced language consisting of different dialects written in various scripts. Approximately 30 million people in different countries speak the language. The lack of corpora is one of the main obstacle…

Researching Less-Resourced Languages -- the DigiSami Corpus

2018-05-01 · LREC 2018 5 · Kristiina Jokinen

Deep Cross-Lingual Coreference Resolution for Less-Resourced Languages: The Case of Basque

2019-06-01 · WS 2019 6 · Gorka Urbizu, Ander Soraluze, Olatz Arregi

In this paper, we present a cross-lingual neural coreference resolution system for a less-resourced language such as Basque. To begin with, we build the first neural coreference resolution system for Basque, training it …

coreference-resolutionCoreference Resolution

Latin-Spanish Neural Machine Translation: from the Bible to Saint Augustine

2020-05-01 · LREC 2020 5 · Eva Mart{\'\i}nez Garcia, {\'A}lvaro Garc{\'\i}a Tejedor

Although there are several sources where to find historical texts, they usually are available in the original language that makes them generally inaccessible. This paper presents the development of state-of-the-art Neura…

Domain AdaptationMachine TranslationTranslation