Developing the Old Tibetan Treebank
This paper presents a full procedure for the development of a segmented, POS-tagged and chunkparsed corpus of Old Tibetan. As an extremely low-resource language, Old Tibetan poses non-trivial problems in every step towards the development of a searchable treebank. We demonstrate, however, that a carefully developed, semisupervised method of optimising and extending existing tools for Classical Tibetan, as well as creating specific ones for Old Tibetan can address these issues. We thus also present the first very Tibetan Treebank in a variety of formats to facilitate research in the fields of NLP, historical linguistics and Tibetan Studies.
Code (0)
등록된 구현이 없습니다.
Tasks
POSSimilar Papers 제목 키워드 기반
TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs
Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions and provide limited coverage of traditi…
Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
Tibetan, one of the major low-resource languages in Asia, presents unique linguistic and sociocultural characteristics that pose both challenges and opportunities for AI research. Despite increasing interest in developin…
Cross-Lingual TransferMachine TranslationSpeech RecognitionTibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation
Tibetan text-to-speech (TTS) has long been challenged by scarce speech resources, significant dialectal variation, and the complex mapping between written text and spoken pronunciation. To address these issues, this work…
Speech SynthesisSun-Shine: A Large Language Model for Tibetan Culture
Tibetan, a minority language in China, features a highly intricate grammatical structure, characterized by four verb tenses and a tense system with frequent irregularities, contributing to its extensive inflectional dive…
Language ModelingLanguage ModellingLarge Language ModelMachine Translation+2A Aelf-supervised Tibetan-chinese Vocabulary Alignment Method Based On Adversarial Learning
Tibetan is a low-resource language. In order to alleviate the shortage of parallel corpus between Tibetan and Chinese, this paper uses two monolingual corpora and a small number of seed dictionaries to learn the semi-sup…