Twitter Language Identification Of Similar Languages And Dialects Without Ground Truth
We present a new method to bootstrap filter Twitter language ID labels in our dataset for automatic language identification (LID). Our method combines geo-location, original Twitter LID labels, and Amazon Mechanical Turk to resolve missing and unreliable labels. We are the first to compare LID classification performance using the MIRA algorithm and langid.py. We show classifier performance on different versions of our dataset with high accuracy using only Twitter data, without ground truth, and very few training examples. We also show how Platt Scaling can be use to calibrate MIRA classifier output values into a probability distribution over candidate classes, making the output more intuitive. Our method allows for fine-grained distinctions between similar languages and dialects and allows us to rediscover the language composition of our Twitter dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationLanguage IdentificationSentiment AnalysisSimilar Papers 제목 키워드 기반
Detection of Similar Languages and Dialects Using Deep Supervised Autoencoder
Language detection is considered a difficult task especially for similar languages, varieties, and dialects. With the growing number of online content in different languages, the need for reliable and robust language det…
Language IdentificationArabic Dialect Identification Using BERT-Based Domain Adaptation
Arabic is one of the most important and growing languages in the world. With the rise of social media platforms such as Twitter, Arabic spoken dialects have become more in use. In this paper, we describe our approach on …
Dialect IdentificationDomain AdaptationDiscrimination between Similar Languages, Varieties and Dialects using CNN- and LSTM-based Deep Neural Networks
In this paper, we describe a system (CGLI) for discriminating similar languages, varieties and dialects using convolutional neural networks (CNNs) and long short-term memory (LSTM) neural networks. We have participated i…
Dialect IdentificationInformation RetrievalLanguage IdentificationMachine Translation+2Findings of the VarDial Evaluation Campaign 2022
This report presents the results of the shared tasks organized as part of the VarDial Evaluation Campaign 2022. The campaign is part of the ninth workshop on Natural Language Processing (NLP) for Similar Languages, Varie…
Dialect IdentificationExtractive Question-AnsweringQuestion AnsweringThe Curious Case of Logistic Regression for Italian Languages and Dialects Identification
Automatic Language Identification represents an important task for improving many real-world applications such as opinion mining and machine translation. In the case of closely-related languages such as regional dialects…
Language IdentificationMachine TranslationOpinion Miningregression+1