paper-with-me

Papers

Developing Language Resources and NLP Tools for the North Korean Language

2022-06-01 · LREC 2022 6 · Arda Akdemir, Yeojoo Jeon, Tetsuo Shibuya

Since the division of Korea, the two Korean languages have diverged significantly over the last 70 years. However, due to the lack of linguistic source of the North Korean language, there is no DPRK-based language model. Consequently, scholars rely on the Korean language model by utilizing South Korean linguistic data. In this paper, we first present a large-scale dataset for the North Korean language. We use the dataset to train a BERT-based language model, DPRK-BERT. Second, we annotate a subset of this dataset for the sentiment analysis task. Finally, we compare the performance of different language models for masked language modeling and sentiment analysis tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingMasked Language ModelingSentiment Analysis

Similar Papers 제목 키워드 기반

Learning How to Translate North Korean through South Korean

2022-01-27 · LREC 2022 6 · Hwichan Kim, Sangwhan Moon, Naoaki Okazaki, Mamoru Komachi

South and North Korea both use the Korean language. However, Korean NLP research has focused on South Korean only, and existing NLP systems of the Korean language, such as neural machine translation (NMT) models, cannot …

Machine TranslationNMTTranslation

Mergen: The First Manchu-Korean Machine Translation Model Trained on Augmented Data

2023-11-29 · Jean Seo, Sungjoo Byun, Minha Kang, Sangah Lee

The Manchu language, with its roots in the historical Manchurian region of Northeast China, is now facing a critical threat of extinction, as there are very few speakers left. In our efforts to safeguard the Manchu langu…

DecoderMachine TranslationTranslation

Zero-shot North Korean to English Neural Machine Translation by Character Tokenization and Phoneme Decomposition

2020-07-01 · ACL 2020 6 · Hwichan Kim, Tosho Hirasawa, Mamoru Komachi

The primary limitation of North Korean to English translation is the lack of a parallel corpus; therefore, high translation accuracy cannot be achieved. To address this problem, we propose a zero-shot approach using Sout…

Machine TranslationTranslation

Why North Korean Refugees are Reluctant to Compete: The Roles of Cognitive Ability

2021-08-18 · Syngjoo Choi, Byung-Yeon Kim, Jungmin Lee, Sokbae Lee

The study compares the competitiveness of three Korean groups raised in different institutional environments: South Korea, North Korea, and China. Laboratory experiments reveal that North Korean refugees are less likely …

Reusing a Multi-lingual Setup to Bootstrap a Grammar Checker for a Very Low Resource Language without Data

2022-05-01 · ComputEL (ACL) 2022 5 · Inga Lill Sigga Mikkelsen, Linda Wiechetek, Flammie A Pirinen

Grammar checkers (GEC) are needed for digital language survival. Very low resource languages like Lule Sámi with less than 3,000 speakers need to hurry to build these tools, but do not have the big corpus data that are r…