Subcharacter Information in Japanese Embeddings: When Is It Worth It?
Languages with logographic writing systems present a difficulty for traditional character-level models. Leveraging the subcharacter information was recently shown to be beneficial for a number of intrinsic and extrinsic tasks in Chinese. We examine whether the same strategies could be applied for Japanese, and contribute a new analogy dataset for this language.
Code (0)
등록된 구현이 없습니다.
Tasks
Text ClassificationSimilar Papers 제목 키워드 기반
New Perspectives in Sinographic Language Processing Through the Use of Character Structure
Chinese characters have a complex and hierarchical graphical structure carrying both semantic and phonetic information. We use this structure to enhance the text model and obtain better results in standard NLP operations…
Relationtext-classificationText ClassificationJoint Embeddings of Chinese Words, Characters, and Fine-grained Subcharacter Components
Word embeddings have attracted much attention recently. Different from alphabetic writing systems, Chinese characters are often composed of subcharacter components which are also semantically informative. In this work, w…
Named Entity Recognition (NER)Question AnsweringSentiment AnalysisText Classification+2SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
Korean is a morphologically rich language with a featural writing system in which each character is systematically composed of subcharacter units known as Jamo. These subcharacters not only determine the visual structure…
Natural Language UnderstandingSpeech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
SpeechLLMs typically combine ASR-trained encoders with text-based LLM backbones, leading them to inherit written-style output patterns unsuitable for text-to-speech synthesis. This mismatch is particularly pronounced in …
Text-To-Speech SynthesisUtilizing Visual Forms of Japanese Characters for Neural Review Classification
We propose a novel method that exploits visual information of ideograms and logograms in analyzing Japanese review documents. Our method first converts font images of Japanese characters into character embeddings using c…
ClassificationGeneral ClassificationSentence