Jamo Pair Encoding: Subcharacter Representation-based Extreme Korean Vocabulary Compression for Efficient Subword Tokenization
In the context of multilingual language model pre-training, vocabulary size for languages with a broad set of potential characters is an unsolved problem. We propose two algorithms applicable in any unsupervised multilingual pre-training task, increasing the elasticity of budget required for building the vocabulary in Byte-Pair Encoding inspired tokenizers, significantly reducing the cost of supporting Korean in a multilingual model.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
Korean is a morphologically rich language with a featural writing system in which each character is systematically composed of subcharacter units known as Jamo. These subcharacters not only determine the visual structure…
Natural Language UnderstandingInvestigating an Effective Character-level Embedding in Korean Sentence Classification
Different from the writing systems of many Romance and Germanic languages, some languages or language families show complex conjunct forms in character composition. For such cases where the conjuncts consist of the compo…
ClassificationGeneral ClassificationSentenceSentence Classification+1Quantifying and Mitigating Korean Jamo-Level Typographical Vulnerabilities in Large Language Models
Korean introduces an additional typographical perturbation level not captured by ordinary character-level edit models: because syllable blocks are internally composed of sub-character units called jamo, keyboard-level er…
Grammatical Error CorrectionJoint Embeddings of Chinese Words, Characters, and Fine-grained Subcharacter Components
Word embeddings have attracted much attention recently. Different from alphabetic writing systems, Chinese characters are often composed of subcharacter components which are also semantically informative. In this work, w…
Named Entity Recognition (NER)Question AnsweringSentiment AnalysisText Classification+2A Sub-Character Architecture for Korean Language Processing
We introduce a novel sub-character architecture that exploits a unique compositional structure of the Korean language. Our method decomposes each character into a small set of primitive phonetic units called jamo letters…
Dependency Parsing