paper-with-me

Papers

Building an Endangered Language Resource in the Classroom: Universal Dependencies for Kakataibo

2022-06-21 · LREC 2022 6 · Roberto Zariquiey, Claudia Alvarado, Ximena Echevarria, Luisa Gomez, Rosa Gonzales, Mariana Illescas, Sabina Oporto, Frederic Blum, Arturo Oncevay, Javier Vera

In this paper, we launch a new Universal Dependencies treebank for an endangered language from Amazonia: Kakataibo, a Panoan language spoken in Peru. We first discuss the collaborative methodology implemented, which proved effective to create a treebank in the context of a Computational Linguistic course for undergraduates. Then, we describe the general details of the treebank and the language-specific considerations implemented for the proposed annotation. We finally conduct some experiments on part-of-speech tagging and syntactic dependency parsing. We focus on monolingual and transfer learning settings, where we study the impact of a Shipibo-Konibo treebank, another Panoan language resource.

📄 PDF Abstract BibTeX arXiv:2206.10343

Code (1)

tarotis/building-an-endangered-language-resource-in-the-classroom 공식 구현 pytorch

Tasks

Dependency ParsingPart-Of-Speech TaggingTransfer Learning

Similar Papers 제목 키워드 기반

When Word Embeddings Become Endangered

2021-03-24 · Khalid Alnajjar

Big languages such as English and Finnish have many natural language processing (NLP) resources and models, but this is not the case for low-resourced and endangered languages as such resources are so scarce despite the …

Cross-Lingual Word EmbeddingsSentiment AnalysisTranslationWord Embeddings

Tusom2021: A Phonetically Transcribed Speech Dataset from an Endangered Language for Universal Phone Recognition Experiments

2021-04-02 · David R. Mortensen, Jordan Picone, Xinjian Li, Kathleen Siminyu

There is growing interest in ASR systems that can recognize phones in a language-independent fashion. There is additionally interest in building language technologies for low-resource and endangered languages. However, t…

Universal Dependency Treebank for Xibe

2020-12-01 · UDW (COLING) 2020 12 · He Zhou, Juyeon Chung, Sandra Kübler, Francis Tyers

We present our work of constructing the first treebank for the Xibe language following the Universal Dependencies (UD) annotation scheme. Xibe is a low-resourced and severely endangered Tungusic language spoken by the Xi…

Recovering Text from Endangered Languages Corrupted PDF documents

2022-05-01 · ComputEL (ACL) 2022 5 · Nicolas Stefanovitch

In this paper we present an approach to efficiently recover texts from corrupted documents of endangered languages. Textual resources for such languages are scarce, and sometimes the few available resources are corrupted…

Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic Information

2024-04-23 · Chihiro Taguchi, Jefferson Saransig, Dayana Velásquez, David Chiang

This paper presents Killkan, the first dataset for automatic speech recognition (ASR) in the Kichwa language, an indigenous language of Ecuador. Kichwa is an extremely low-resource endangered language, and there have bee…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition