Variation and Instability in Dialect-Based Embedding Spaces
This paper measures variation in embedding spaces which have been trained on different regional varieties of English while controlling for instability in the embeddings. While previous work has shown that it is possible to distinguish between similar varieties of a language, this paper experiments with two follow-up questions: First, does the variety represented in the training data systematically influence the resulting embedding space after training? This paper shows that differences in embeddings across varieties are significantly higher than baseline instability. Second, is such dialect-based variation spread equally throughout the lexicon? This paper shows that specific parts of the lexicon are particularly subject to variation. Taken together, these experiments confirm that embedding spaces are significantly influenced by the dialect represented in the training data. This finding implies that there is semantic variation across dialects, in addition to previously-studied lexical and syntactic variation.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
DialectGram: Detecting Dialectal Variation at Multiple Geographic Resolutions
Several computational models have been developed to detect and analyze dialect variation in recent years. Most of these models assume a predefined set of geographical regions over which they detect and analyze dialectal …
Word EmbeddingsThe Effects of Randomness on the Stability of Node Embeddings
We systematically evaluate the (in-)stability of state-of-the-art node embedding algorithms due to randomness, i.e., the random variation of their outcomes given identical algorithms and graphs. We apply five node embedd…
General ClassificationNode ClassificationModeling Orthographic Variation in Occitan's Dialects
Effectively normalizing textual data poses a considerable challenge, especially for low-resource languages lacking standardized writing systems. In this study, we fine-tuned a multilingual model with data from several Oc…
Dependency ParsingPart-Of-Speech TaggingTowards One Model to Rule All: Multilingual Strategy for Dialectal Code-Switching Arabic ASR
With the advent of globalization, there is an increasing demand for multilingual automatic speech recognition (ASR), handling language and dialectal variation of spoken content. Recent studies show its efficacy over mono…
AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1Dialectograms: Machine Learning Differences between Discursive Communities
Word embeddings provide an unsupervised way to understand differences in word usage between discursive communities. A number of recent papers have focused on identifying words that are used differently by two or more com…
Word Embeddings