Addressing Low-Resource Scenarios with Character-aware Embeddings
Most modern approaches to computing word embeddings assume the availability of text corpora with billions of words. In this paper, we explore a setup where only corpora with millions of words are available, and many words in any new text are out of vocabulary. This setup is both of practical interests {--} modeling the situation for specific domains and low-resource languages {--} and of psycholinguistic interest, since it corresponds much more closely to the actual experiences and challenges of human language learning and use. We compare standard skip-gram word embeddings with character-based embeddings on word relatedness prediction. Skip-grams excel on large corpora, while character-based embeddings do well on small corpora generally and rare and complex words specifically. The models can be combined easily.
Code (0)
등록된 구현이 없습니다.
Tasks
Morphological AnalysisWord EmbeddingsSimilar Papers 제목 키워드 기반
Multiple Object Tracking based on Occlusion-Aware Embedding Consistency Learning
The Joint Detection and Embedding (JDE) framework has achieved remarkable progress for multiple object tracking. Existing methods often employ extracted embeddings to re-establish associations between new detections and …
Multiple Object TrackingObjectObject TrackingvalidLAPS-Diff: A Diffusion-Based Framework for Singing Voice Synthesis With Language Aware Prosody-Style Guided Learning
The field of Singing Voice Synthesis (SVS) has seen significant advancements in recent years due to the rapid progress of diffusion-based approaches. However, capturing vocal style, genre-specific pitch inflections, and …
MultiSeg: Parallel Data and Subword Information for Learning Bilingual Embeddings in Low Resource Scenarios
Distributed word embeddings have become ubiquitous in natural language processing as they have been shown to improve performance in many semantic and syntactic tasks. Popular models for learning cross-lingual word embedd…
Cross-Lingual Word EmbeddingsTranslationWord EmbeddingsWord Similarity+1RL-Driven Security-Aware Resource Allocation Framework for UAV-Assisted O-RAN
The integration of Unmanned Aerial Vehicles (UAVs) into Open Radio Access Networks (O-RAN) enhances communication in disaster management and Search and Rescue (SAR) operations by ensuring connectivity when infrastructure…
Reinforcement LearningIntroducing Syllable Tokenization for Low-resource Languages: A Case Study with Swahili
Many attempts have been made in multilingual NLP to ensure that pre-trained language models, such as mBERT or GPT2 get better and become applicable to low-resource languages. To achieve multilingualism for pre-trained la…
Multilingual NLPText GenerationWord Embeddings