Learning to Generate Word Representations using Subword Information
Distributed representations of words play a major role in the field of natural language processing by encoding semantic and syntactic information of words. However, most existing works on learning word representations typically regard words as individual atomic units and thus are blind to subword information in words. This further gives rise to a difficulty in representing out-of-vocabulary (OOV) words. In this paper, we present a character-based word representation approach to deal with this limitation. The proposed model learns to generate word representations from characters. In our model, we employ a convolutional neural network and a highway network over characters to extract salient features effectively. Unlike previous models that learn word representations from a large corpus, we take a set of pre-trained word embeddings and generalize it to word entries, including OOV words. To demonstrate the efficacy of the proposed model, we perform both an intrinsic and an extrinsic task which are word similarity and language modeling, respectively. Experimental results show clearly that the proposed model significantly outperforms strong baseline models that regard words or their subwords as atomic units. For example, we achieve as much as 18.5{\%} improvement on average in perplexity for morphologically rich languages compared to strong baselines in the language modeling task.
Code (0)
등록된 구현이 없습니다.
Tasks
ChunkingLanguage ModelingLanguage ModellingNamed Entity Recognition (NER)Question AnsweringText ClassificationWord EmbeddingsWord SimilarityMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Systematic Study of Leveraging Subword Information for Learning Word Representations
The use of subword-level information (e.g., characters, character n-grams, morphemes) has become ubiquitous in modern word representation learning. Its importance is attested especially for morphologically rich languages…
Dependency ParsingEntity TypingRepresentation LearningSegmentation+1Learning Mutually Informed Representations for Characters and Subwords
Most pretrained language models rely on subword tokenization, which processes text as a sequence of subword tokens. However, different granularities of text, such as characters, subwords, and words, can contain different…
named-entity-recognitionNamed Entity RecognitionPOSPOS Tagging+2Incorporating Subword Information into Matrix Factorization Word Embeddings
The positive effect of adding subword information to word embeddings has been demonstrated for predictive models. In this paper we investigate whether similar benefits can also be derived from incorporating subwords into…
Word EmbeddingsAudio-to-Intent Using Acoustic-Textual Subword Representations from End-to-End ASR
Accurate prediction of the user intent to interact with a voice assistant (VA) on a device (e.g. on the phone) is critical for achieving naturalistic, engaging, and privacy-centric interactions with the VA. To this end, …
intent-classificationIntent ClassificationCombining Subword Representations into Word-level Representations in the Transformer Architecture
In Neural Machine Translation, using word-level tokens leads to degradation in translation quality. The dominant approaches use subword-level tokens, but this increases the length of the sequences and makes it difficult …
Machine TranslationPOSTranslation