A Dependency Treebank of the Chinese Buddhist Canon
We present a dependency treebank of the Chinese Buddhist Canon, which contains 1,514 texts with about 50 million Chinese characters. The treebank was created by an automatic parser trained on a smaller treebank, containing four manually annotated sutras (Lee and Kong, 2014). We report results on word segmentation, part-of-speech tagging and dependency parsing, and discuss challenges posed by the processing of medieval Chinese. In a case study, we exploit the treebank to examine verbs frequently associated with Buddha, and to analyze usage patterns of quotative verbs in direct speech. Our results suggest that certain quotative verbs imply status differences between the speaker and the listener.
Code (0)
등록된 구현이 없습니다.
Tasks
Dependency ParsingPart-Of-Speech TaggingSimilar Papers 제목 키워드 기반
Converting the Sinica Treebank of Mandarin Chinese to Universal Dependencies
This paper describes the conversion of the Sinica Treebank, one of the major Mandarin Chinese treebanks, to Universal Dependencies. The conversion is rule-based and the process involves POS tag mapping, head adjusting in…
POSTAGDependency Network Syntax: From Dependency Treebanks to a Classification of Chinese Function Words
Developing Universal Dependencies for Mandarin Chinese
This article proposes a Universal Dependency Annotation Scheme for Mandarin Chinese, including POS tags and dependency analysis. We identify cases of idiosyncrasy of Mandarin Chinese that are difficult to fit into the cu…
POSBlur the Linguistic Boundary: Interpreting Chinese Buddhist Sutra in English via Neural Machine Translation
Buddhism is an influential religion with a long-standing history and profound philosophy. Nowadays, more and more people worldwide aspire to learn the essence of Buddhism, attaching importance to Buddhism dissemination. …
Machine TranslationNMTPhilosophyTranslation