Addressing Noise in Multidialectal Word Embeddings
Word embeddings are crucial to many natural language processing tasks. The quality of embeddings relies on large non-noisy corpora. Arabic dialects lack large corpora and are noisy, being linguistically disparate with no standardized spelling. We make three contributions to address this noise. First, we describe simple but effective adaptations to word embedding tools to maximize the informative content leveraged in each training sentence. Second, we analyze methods for representing disparate dialects in one embedding space, either by mapping individual dialects into a shared space or learning a joint model of all dialects. Finally, we evaluate via dictionary induction, showing that two metrics not typically reported in the task enable us to analyze our contributions{'} effects on low and high frequency words. In addition to boosting performance between 2-53{\%}, we specifically improve on noisy, low frequency forms without compromising accuracy on high frequency forms.
Code (0)
등록된 구현이 없습니다.
Tasks
SentenceTransliterationWord EmbeddingsSimilar Papers 제목 키워드 기반
DiaLex: A Benchmark for Evaluating Multidialectal Arabic Word Embeddings
Word embeddings are a core component of modern natural language processing systems, making the ability to thoroughly evaluate them a vital task. We describe DiaLex, a benchmark for intrinsic evaluation of dialectal Arabi…
Word EmbeddingsProMap: Effective Bilingual Lexicon Induction via Language Model Prompting
Bilingual Lexicon Induction (BLI), where words are translated between two languages, is an important NLP task. While noticeable progress on BLI in rich resource languages using static word embeddings has been achieved. T…
Bilingual Lexicon InductionLanguage ModelingLanguage ModellingRe-Ranking+3Crowdsourcing Latin American Spanish for Low-Resource Text-to-Speech
In this paper we present a multidialectal corpus approach for building a text-to-speech voice for a new dialect in a language with existing resources, focusing on various South American dialects of Spanish. We first pres…
text-to-speechText to SpeechNeural-based Noise Filtering from Word Embeddings
Word embeddings have been demonstrated to benefit NLP tasks impressively. Yet, there is room for improvement in the vector representations, because current word embeddings typically contain unnecessary information, i.e.,…
DenoisingWord EmbeddingsOn Learning Word Embeddings From Linguistically Augmented Text Corpora
Word embedding is a technique in Natural Language Processing (NLP) to map words into vector space representations. Since it has boosted the performance of many NLP downstream tasks, the task of learning word embeddings h…
Learning Word EmbeddingsWord EmbeddingsWord Similarity