paper-with-me

홈 › Papers

Wikinflection Corpus: A (Better) Multilingual, Morpheme-Annotated Inflectional Corpus

2020-05-01 · LREC 2020 5 · Eleni Metheniti, Guenter Neumann

Multilingual, inflectional corpora are a scarce resource in the NLP community, especially corpora with annotated morpheme boundaries. We are evaluating a generated, multilingual inflectional corpus with morpheme boundaries, generated from the English Wiktionary (Metheniti and Neumann, 2018), against the largest, multilingual, high-quality inflectional corpus of the UniMorph project (Kirov et al., 2018). We confirm that the generated Wikinflection corpus is not of such quality as UniMorph, but we were able to extract a significant amount of words from the intersection of the two corpora. Our Wikinflection corpus benefits from the morpheme segmentations of Wiktionary/Wikinflection and from the manually-evaluated morphological feature tags of the UniMorph project, and has 216K lemmas and 5.4M word forms, in a total of 68 languages.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text

2024-03-11 · Michael Ginn, Lindia Tjuatja, Taiqi He, Enora Rice 외

Language documentation projects often involve the creation of annotated text in a format such as interlinear glossed text (IGT), which captures fine-grained morphosyntactic analyses in a morpheme-by-morpheme format. Howe…

Expanding Universal Dependencies for Polysynthetic Languages: A Case of St. Lawrence Island Yupik

2021-06-01 · NAACL (AmericasNLP) 2021 6 · Hyunji Park, Lane Schwartz, Francis Tyers

This paper describes the development of the first Universal Dependencies (UD) treebank for St. Lawrence Island Yupik, an endangered language spoken in the Bering Strait region. While the UD guidelines provided a general …

Dependency Parsing

A Sentiment Corpus for South African Under-Resourced Languages in a Multilingual Context

2022-06-01 · SIGUL (LREC) 2022 6 · Ronny Mabokela, Tim Schlippe

Multilingual sentiment analysis is a process of detecting and classifying sentiment based on textual information written in multiple languages. There has been tremendous research advancement on high-resourced languages s…

Sentiment Analysis

The making of the Litkey Corpus, a richly annotated longitudinal corpus of German texts written by primary school children

2019-08-01 · WS 2019 8 · Ronja Laarmann-Quante, Stefanie Dipper, Eva Belke

To date, corpus and computational linguistic work on written language acquisition has mostly dealt with second language learners who have usually already mastered orthography acquisition in their first language. In this …

Language AcquisitionPOS

Massively Multilingual Joint Segmentation and Glossing

2026-01-16 · Michael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian 외 arxiv

Automated interlinear gloss prediction with neural networks is a promising approach to accelerate language documentation efforts. However, while state-of-the-art models like GlossLM achieve high scores on glossing benchm…