paper-with-me

홈 › Papers

NLP Pipeline for Annotating (Endangered) Tibetan and Newar Varieties

2022-06-01 · EURALI (LREC) 2022 6 · Christian Faggionato, Nathan Hill, Marieke Meelen

In this paper we present our work-in-progress on a fully-implemented pipeline to create deeply-annotated corpora of a number of historical and contemporary Tibetan and Newar varieties. Our off-the-shelf tools allow researchers to create corpora with five different layers of annotation, ranging from morphosyntactic to information-structural annotation. We build on and optimise existing tools (in line with FAIR principles), as well as develop new ones, and show how they can be adapted to other Tibetan and Newar languages, most notably modern endangered languages that are both extremely low-resourced and under-researched.

📄 PDF Abstract BibTeX

Code (1)

lothelanor/actib 공식 구현

Similar Papers 제목 키워드 기반

Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR

2026-03-08 · Rishikesh Kumar Sharma, Safal Narshing Shrestha, Jenny Poudel, Rupak Tiwari 외 arxiv

Nepal Bhasha (Newari), an endangered language of the Kathmandu Valley, remains digitally marginalized due to the severe scarcity of annotated speech resources. In this work, we introduce Nwāchā Munā, a newly curated 5.39…

Cross-Lingual TransferSpeech RecognitionData Augmentation

Automatic Extraction of Verb Paradigms in Regional Languages: the case of the Linguistic Crescent varieties

2020-05-01 · LREC 2020 5 · elena knyazeva, Gilles Adda, Philippe Boula de Mare{\"u}il, Maximilien Gu{\'e}rin 외

Language documentation is crucial for endangered varieties all over the world. Verb conjugation is a key aspect of this documentation for Romance varieties such as those spoken in central France, in the area of the Lingu…

From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for Tibetan

2025-07-12 · Lei Yang, Leiyu Pan, Bojian Xiong, Renren Jin 외 arxiv

Large language models (LLMs) have achieved remarkable success across a wide range of natural language processing tasks, yet their performance remains heavily biased toward high-resource languages. Tibetan, despite its cu…

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling

2026-05-26 · Guixian Xu, Yide Liang, Zeli Su, Xuexian Song 외 arxiv

Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible training and evaluation infrastructure. To fill this gap, we introduce FTib…

Continual Pretraining

The Tonogenesis Continuum in Tibetan: A Computational Investigation

2025-10-26 · Siyu Liang, Zhaxi Zerong arxiv

Tonogenesis-the historical process by which segmental contrasts evolve into lexical tone-has traditionally been studied through comparative reconstruction and acoustic phonetics. We introduce a computational approach tha…

Speech Recognition