paper-with-me

홈 › Papers

Twitter Language Identification Of Similar Languages And Dialects Without Ground Truth

2017-04-01 · WS 2017 4 · Jennifer Williams, Charlie Dagli

We present a new method to bootstrap filter Twitter language ID labels in our dataset for automatic language identification (LID). Our method combines geo-location, original Twitter LID labels, and Amazon Mechanical Turk to resolve missing and unreliable labels. We are the first to compare LID classification performance using the MIRA algorithm and langid.py. We show classifier performance on different versions of our dataset with high accuracy using only Twitter data, without ground truth, and very few training examples. We also show how Platt Scaling can be use to calibrate MIRA classifier output values into a probability distribution over candidate classes, making the output more intuitive. Our method allows for fine-grained distinctions between similar languages and dialects and allows us to rediscover the language composition of our Twitter dataset.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

General ClassificationLanguage IdentificationSentiment Analysis

Similar Papers 제목 키워드 기반

Detection of Similar Languages and Dialects Using Deep Supervised Autoencoder

2020-12-01 · ICON 2020 12 · Shantipriya Parida, Esau Villatoro-Tello, Sajit Kumar, Maël Fabien 외

Language detection is considered a difficult task especially for similar languages, varieties, and dialects. With the growing number of online content in different languages, the need for reliable and robust language det…

Language Identification

Arabic Dialect Identification Using BERT-Based Domain Adaptation

2020-11-13 · COLING (WANLP) 2020 12 · Ahmad Beltagy, Abdelrahman Wael, Omar ElSherief

Arabic is one of the most important and growing languages in the world. With the rise of social media platforms such as Twitter, Arabic spoken dialects have become more in use. In this paper, we describe our approach on …

Dialect IdentificationDomain Adaptation

Discrimination between Similar Languages, Varieties and Dialects using CNN- and LSTM-based Deep Neural Networks

2016-12-01 · WS 2016 12 · Chinnappa Guggilla

In this paper, we describe a system (CGLI) for discriminating similar languages, varieties and dialects using convolutional neural networks (CNNs) and long short-term memory (LSTM) neural networks. We have participated i…

Dialect IdentificationInformation RetrievalLanguage IdentificationMachine Translation+2

Findings of the VarDial Evaluation Campaign 2022

2022-10-01 · VarDial (COLING) 2022 10 · Noëmi Aepli, Antonios Anastasopoulos, Adrian-Gabriel Chifu, William Domingues 외

This report presents the results of the shared tasks organized as part of the VarDial Evaluation Campaign 2022. The campaign is part of the ninth workshop on Natural Language Processing (NLP) for Similar Languages, Varie…

Dialect IdentificationExtractive Question-AnsweringQuestion Answering

The Curious Case of Logistic Regression for Italian Languages and Dialects Identification

2022-10-01 · VarDial (COLING) 2022 10 · Giacomo Camposampiero, Quynh Anh Nguyen, Francesco Di Stefano

Automatic Language Identification represents an important task for improving many real-world applications such as opinion mining and machine translation. In the case of closely-related languages such as regional dialects…

Language IdentificationMachine TranslationOpinion Miningregression+1