paper-with-me

Papers

Developing the Old Tibetan Treebank

2019-09-01 · RANLP 2019 9 · Christian Faggionato, Marieke Meelen

This paper presents a full procedure for the development of a segmented, POS-tagged and chunkparsed corpus of Old Tibetan. As an extremely low-resource language, Old Tibetan poses non-trivial problems in every step towards the development of a searchable treebank. We demonstrate, however, that a carefully developed, semisupervised method of optimising and extending existing tools for Classical Tibetan, as well as creating specific ones for Old Tibetan can address these issues. We thus also present the first very Tibetan Treebank in a variety of formats to facilitate research in the fields of NLP, historical linguistics and Tibetan Studies.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

POS

Similar Papers 제목 키워드 기반

TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs

2026-08-01 · Jin Zhang, Linyu Li, Weili Jiang, Yuqing Cai 외 arxiv

Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions and provide limited coverage of traditi…

Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges

2025-10-22 · Cheng Huang, Nyima Tashi, Fan Gao, Yutong Liu 외 arxiv

Tibetan, one of the major low-resource languages in Asia, presents unique linguistic and sociocultural characteristics that pose both challenges and opportunities for AI research. Despite increasing interest in developin…

Cross-Lingual TransferMachine TranslationSpeech Recognition

Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation

2026-05-04 · Jiaxu He, Chao Wang, Jie Lian, Yuqing Cai 외 arxiv

Tibetan text-to-speech (TTS) has long been challenged by scarce speech resources, significant dialectal variation, and the complex mapping between written text and spoken pronunciation. To address these issues, this work…

Speech Synthesis

Sun-Shine: A Large Language Model for Tibetan Culture

2025-03-24 · Cheng Huang, Fan Gao, Nyima Tashi, Yutong Liu 외

Tibetan, a minority language in China, features a highly intricate grammatical structure, characterized by four verb tenses and a tense system with frequent irregularities, contributing to its extensive inflectional dive…

Language ModelingLanguage ModellingLarge Language ModelMachine Translation+2

A Aelf-supervised Tibetan-chinese Vocabulary Alignment Method Based On Adversarial Learning

2021-10-04 · Enshuai Hou, Jie Zhu

Tibetan is a low-resource language. In order to alleviate the shortage of parallel corpus between Tibetan and Chinese, this paper uses two monolingual corpora and a small number of seed dictionaries to learn the semi-sup…