Researching Less-Resourced Languages -- the DigiSami Corpus
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
EM Corpus: a comparable corpus for a less-resourced language pair Manipuri-English
In this paper, we introduce a sentence-level comparable text corpus crawled and created for the less-resourced language pair, Manipuri(mni) and English (eng). Our monolingual corpora comprise 1.88 million Manipuri senten…
SentenceLR-Sum: Summarization for Less-Resourced Languages
This preprint describes work in progress on LR-Sum, a new permissively-licensed dataset created with the goal of enabling further research in automatic summarization for less-resourced languages. LR-Sum contains human-wr…
Building a Bilingual Vietnamese-French Named Entity Annotated Corpus through Cross-Linguistic Projection
The creation of high-quality named entity annotated resources is time-consuming and an expensive process. Most of the gold standard corpora are available for English but not for less-resourced languages such as Vietnames…
Deep Cross-Lingual Coreference Resolution for Less-Resourced Languages: The Case of Basque
In this paper, we present a cross-lingual neural coreference resolution system for a less-resourced language such as Basque. To begin with, we build the first neural coreference resolution system for Basque, training it …
coreference-resolutionCoreference ResolutionCrawling Under-Resourced Languages - a Portal for Community-Contributed Corpus Collection
The “Web as corpus” paradigm opens opportunities for enhancing the current state of language resources for endangered and under-resourced languages. However, standard crawling strategies tend to overlook available resour…