Developing Language Resources and NLP Tools for the North Korean Language
Since the division of Korea, the two Korean languages have diverged significantly over the last 70 years. However, due to the lack of linguistic source of the North Korean language, there is no DPRK-based language model. Consequently, scholars rely on the Korean language model by utilizing South Korean linguistic data. In this paper, we first present a large-scale dataset for the North Korean language. We use the dataset to train a BERT-based language model, DPRK-BERT. Second, we annotate a subset of this dataset for the sentiment analysis task. Finally, we compare the performance of different language models for masked language modeling and sentiment analysis tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingMasked Language ModelingSentiment AnalysisSimilar Papers 제목 키워드 기반
Learning How to Translate North Korean through South Korean
South and North Korea both use the Korean language. However, Korean NLP research has focused on South Korean only, and existing NLP systems of the Korean language, such as neural machine translation (NMT) models, cannot …
Machine TranslationNMTTranslationMergen: The First Manchu-Korean Machine Translation Model Trained on Augmented Data
The Manchu language, with its roots in the historical Manchurian region of Northeast China, is now facing a critical threat of extinction, as there are very few speakers left. In our efforts to safeguard the Manchu langu…
DecoderMachine TranslationTranslationZero-shot North Korean to English Neural Machine Translation by Character Tokenization and Phoneme Decomposition
The primary limitation of North Korean to English translation is the lack of a parallel corpus; therefore, high translation accuracy cannot be achieved. To address this problem, we propose a zero-shot approach using Sout…
Machine TranslationTranslationWhy North Korean Refugees are Reluctant to Compete: The Roles of Cognitive Ability
The study compares the competitiveness of three Korean groups raised in different institutional environments: South Korea, North Korea, and China. Laboratory experiments reveal that North Korean refugees are less likely …
Reusing a Multi-lingual Setup to Bootstrap a Grammar Checker for a Very Low Resource Language without Data
Grammar checkers (GEC) are needed for digital language survival. Very low resource languages like Lule Sámi with less than 3,000 speakers need to hurry to build these tools, but do not have the big corpus data that are r…