paper-with-me

홈 › Papers

Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning

2025-07-05 · Nayeon Kim, Eojin Jeon, Jun-Hyung Park, SangKeun Lee arxiv

In this study, we introduce KOPL, a novel framework for handling Korean OOV words with Phoneme representation Learning. Our work is based on the linguistic property of Korean as a phonemic script, the high correlation between phonemes and letters. KOPL incorporates phoneme and word representations for Korean OOV words, facilitating Korean OOV word representations to capture both text and phoneme information of words. We empirically demonstrate that KOPL significantly improves the performance on Korean Natural Language Processing (NLP) tasks, while being readily integrated into existing static and contextual Korean embedding models in a plug-and-play manner. Notably, we show that KOPL outperforms the state-of-the-art model by an average of 1.9%. Our code is available at https://github.com/jej127/KOPL.git.

📄 PDF Abstract BibTeX arXiv:2507.04018

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

A Syllable-based Technique for Word Embeddings of Korean Words

2017-08-05 · WS 2017 9 · Sanghyuk Choi, Taeuk Kim, Jinseok Seol, Sang-goo Lee

Word embedding has become a fundamental component to many NLP tasks such as named entity recognition and machine translation. However, popular models that learn such embeddings are unaware of the morphology of words, so …

Machine Translationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2

Korean-to-Chinese Machine Translation using Chinese Character as Pivot Clue

2019-11-25 · Jeonghyeok Park, Hai Zhao

Korean-Chinese is a low resource language pair, but Korean and Chinese have a lot in common in terms of vocabulary. Sino-Korean words, which can be converted into corresponding Chinese characters, account for more than f…

Machine TranslationTranslation

EVOKE: Emotion Vocabulary Of Korean and English

2026-02-11 · Yoonwon Jung, Hagyeong Shin, Benjamin K. Bergen arxiv

This paper introduces EVOKE (Emotion Vocabulary of Korean and English), a Korean-English parallel dataset of emotion words. The dataset offers comprehensive coverage of emotion words in each language, in addition to many…

Korean-to-Japanese Neural Machine Translation System using Hanja Information

2020-12-01 · AACL (WAT) 2020 12 · Hwichan Kim, Tosho Hirasawa, Mamoru Komachi

In this paper, we describe our TMU neural machine translation (NMT) system submitted for the Patent task (Korean→Japanese) of the 7th Workshop on Asian Translation (WAT 2020, Nakazawa et al., 2020). We propose a novel me…

Machine TranslationNMTTranslation

Korean L2 Vocabulary Prediction: Can a Large Annotated Corpus be Used to Train Better Models for Predicting Unknown Words?

2018-05-01 · LREC 2018 5 · Kevin Yancey, Yves Lepage
Language AcquisitionLexical SimplificationReading ComprehensionText Simplification