paper-with-me

홈 › Papers

A Resource for Computational Experiments on Mapudungun

2019-12-04 · LREC 2020 5 · Mingjun Duan, Carlos Fasola, Sai Krishna Rallabandi, Rodolfo M. Vega, Antonios Anastasopoulos, Lori Levin, Alan W. black

We present a resource for computational experiments on Mapudungun, a polysynthetic indigenous language spoken in Chile with upwards of 200 thousand speakers. We provide 142 hours of culturally significant conversations in the domain of medical treatment. The conversations are fully transcribed and translated into Spanish. The transcriptions also include annotations for code-switching and non-standard pronunciations. We also provide baseline results on three core NLP tasks: speech recognition, speech synthesis, and machine translation between Spanish and Mapudungun. We further explore other applications for which the corpus will be suitable, including the study of code-switching, historical orthography change, linguistic structure, and sociological and anthropological studies.

📄 PDF Abstract BibTeX arXiv:1912.01772

Code (1)

mingjund/mapudungun-corpus 공식 구현

Tasks

Machine Translationspeech-recognitionSpeech RecognitionSpeech SynthesisTranslation

Similar Papers 제목 키워드 기반

Valency Classification of Mapudungun Verbal Roots. Established by the language's own morphotactics

2026-04-01 · Andrés Chandía arxiv

In the previous work, a lexical (re)categorisation -- or confirmation of the given category -- of roots identified as verbal was undertaken to determine their original category accurately. Building on this, the present p…

Adapting to the Low-Resource Double-Bind: Investigating Low-Compute Methods on Low-Resource African Languages

2023-03-29 · Colin Leong, Herumb Shandilya, Bonaventure F. P. Dossou, Atnafu Lambebo Tonja 외

Many natural language processing (NLP) tasks make use of massively pre-trained language models, which are computationally expensive. However, access to high computational resources added to the issue of data scarcity of …

A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments

2017-10-10 · LREC 2018 5 · P. Godard, G. Adda, M. Adda-Decker, J. Benjumea 외

Most speech and language technologies are trained with massive amounts of speech and text information. However, most of the world languages do not have such resources or stable orthography. Systems constructed under thes…

A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments

2018-05-01 · LREC 2018 5 · Pierre Godard, Gilles Adda, Martine Adda-Decker, Juan Benjumea 외

SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time Scaling

2025-11-29 · Yang Xiao, Chunpu Xu, Ruifeng Yuan, Jiashuo Wang 외 arxiv

Test-time compute scaling has emerged as a powerful paradigm for enhancing mathematical reasoning in large language models (LLMs) by allocating additional computational resources during inference. However, current method…

Mathematical Reasoning