UniDic for Early Middle Japanese: a Dictionary for Morphological Analysis of Classical Japanese
In order to construct an annotated diachronic corpus of Japanese, we propose to create a new dictionary for morphological analysis of Early Middle Japanese (Classical Japanese) based on UniDic, a dictionary for Contemporary Japanese. Differences between the Early Middle Japanese and Contemporary Japanese, which prevent a na{\"\i}ve adaptation of UniDic to Early Middle Japanese, are found at the levels of lexicon, morphology, grammar, orthography and pronunciation. In order to overcome these problems, we extended dictionary entries and created a training corpus of Early Middle Japanese to adapt UniDic for Contemporary Japanese to Early Middle Japanese. Experimental results show that the proposed UniDic-EMJ, a new dictionary for Early Middle Japanese, achieves as high accuracy (97{\%}) as needed for the linguistic research on lexicon and grammar in Japanese classical text analysis.
Code (0)
등록된 구현이 없습니다.
Tasks
Morphological AnalysisSimilar Papers 제목 키워드 기반
Accent Estimation of Japanese Words from Their Surfaces and Romanizations for Building Large Vocabulary Accent Dictionaries
In Japanese text-to-speech (TTS), it is necessary to add accent information to the input sentence. However, there are a limited number of publicly available accent dictionaries, and those dictionaries e.g. UniDic, do not…
Sentencetext-to-speechText to SpeechDictionary Look-up with Katakana Variant Recognition
The Japanese language has rich variety and quantity of word variant. Since 1980s, it has been recognized that this richness becomes an obstacle against natural language processing. A complete solution, however, has not b…
Morphological AnalysisRetrievalTransliterationUniversal Dependencies for Japanese
We present an attempt to port the international syntactic annotation scheme, Universal Dependencies, to the Japanese language in this paper. Since the Japanese syntactic structure is usually annotated on the basis of uni…
Shrinking Japanese Morphological Analyzers With Neural Networks and Semi-supervised Learning
For languages without natural word boundaries, like Japanese and Chinese, word segmentation is a prerequisite for downstream analysis. For Japanese, segmentation is often done jointly with part of speech tagging, and thi…
Chinese Word SegmentationMorphological AnalysisPart-Of-Speech TaggingSegmentationBack to Patterns: Efficient Japanese Morphological Analysis with Feature-Sequence Trie
Accurate neural models are much less efficient than non-neural models and are useless for processing billions of social media posts or handling user queries in real time with a limited budget. This study revisits the fas…
CPUMorphological Analysis