Method of Tibetan Person Knowledge Extraction
Person knowledge extraction is the foundation of the Tibetan knowledge graph construction, which provides support for Tibetan question answering system, information retrieval, information extraction and other researches, and promotes national unity and social stability. This paper proposes a SVM and template based approach to Tibetan person knowledge extraction. Through constructing the training corpus, we build the templates based the shallow parsing analysis of Tibetan syntactic, semantic features and verbs. Using the training corpus, we design a hierarchical SVM classifier to realize the entity knowledge extraction. Finally, experimental results prove the method has greater improvement in Tibetan person knowledge extraction.
Code (0)
등록된 구현이 없습니다.
Tasks
graph constructionInformation RetrievalQuestion AnsweringRetrievalUnityMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Sun-Shine: A Large Language Model for Tibetan Culture
Tibetan, a minority language in China, features a highly intricate grammatical structure, characterized by four verb tenses and a tense system with frequent irregularities, contributing to its extensive inflectional dive…
Language ModelingLanguage ModellingLarge Language ModelMachine Translation+2Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation
Tibetan text-to-speech (TTS) has long been challenged by scarce speech resources, significant dialectal variation, and the complex mapping between written text and spoken pronunciation. To address these issues, this work…
Speech SynthesisTreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs
Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions and provide limited coverage of traditi…
Developing the Old Tibetan Treebank
This paper presents a full procedure for the development of a segmented, POS-tagged and chunkparsed corpus of Old Tibetan. As an extremely low-resource language, Old Tibetan poses non-trivial problems in every step towar…
POSA Aelf-supervised Tibetan-chinese Vocabulary Alignment Method Based On Adversarial Learning
Tibetan is a low-resource language. In order to alleviate the shortage of parallel corpus between Tibetan and Chinese, this paper uses two monolingual corpora and a small number of seed dictionaries to learn the semi-sup…