paper-with-me

Papers

Jamo Pair Encoding: Subcharacter Representation-based Extreme Korean Vocabulary Compression for Efficient Subword Tokenization

2020-05-01 · LREC 2020 5 · Sangwhan Moon, Naoaki Okazaki

In the context of multilingual language model pre-training, vocabulary size for languages with a broad set of potential characters is an unsolved problem. We propose two algorithms applicable in any unsupervised multilingual pre-training task, increasing the elasticity of budget required for building the vocabulary in Byte-Pair Encoding inspired tokenizers, significantly reducing the cost of supporting Korean in a multilingual model.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models

2026-04-14 · SungHo Kim, Juhyeong Park, Eda Atalay, SangKeun Lee arxiv

Korean is a morphologically rich language with a featural writing system in which each character is systematically composed of subcharacter units known as Jamo. These subcharacters not only determine the visual structure…

Natural Language Understanding

Investigating an Effective Character-level Embedding in Korean Sentence Classification

2019-05-31 · Won Ik Cho, Seok Min Kim, Nam Soo Kim

Different from the writing systems of many Romance and Germanic languages, some languages or language families show complex conjunct forms in character composition. For such cases where the conjuncts consist of the compo…

ClassificationGeneral ClassificationSentenceSentence Classification+1

Quantifying and Mitigating Korean Jamo-Level Typographical Vulnerabilities in Large Language Models

2026-08-31 · Seojin Lee, Hwanhee Lee arxiv

Korean introduces an additional typographical perturbation level not captured by ordinary character-level edit models: because syllable blocks are internally composed of sub-character units called jamo, keyboard-level er…

Grammatical Error Correction

Joint Embeddings of Chinese Words, Characters, and Fine-grained Subcharacter Components

2017-09-01 · EMNLP 2017 9 · Jinxing Yu, Xun Jian, Hao Xin, Yangqiu Song

Word embeddings have attracted much attention recently. Different from alphabetic writing systems, Chinese characters are often composed of subcharacter components which are also semantically informative. In this work, w…

Named Entity Recognition (NER)Question AnsweringSentiment AnalysisText Classification+2

A Sub-Character Architecture for Korean Language Processing

2017-07-20 · EMNLP 2017 9 · Karl Stratos

We introduce a novel sub-character architecture that exploits a unique compositional structure of the Korean language. Our method decomposes each character into a small set of primitive phonetic units called jamo letters…

Dependency Parsing